8 papers
Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers
Tim Tsz-Kit Lau, Weijie Su
A striking geometric disparity has long persisted in the practice of deep learning. While modern neural network architectures naturally exhibit rich symmetry and equivariance prope…
When Does Model Collapse Occur in Structured Interactive Learning?
Yuchen Wu, Kangjie Zhou, Weijie Su
The proliferation of generative artificial intelligence has given rise to an interactive learning environment, where model parameters are continuously updated using not only data g…
Uncovering Symmetry Transfer in Large Language Models via Layer-Peeled Optimization
Zhehang Du, Hangfeng He, Weijie Su
Large language models (LLMs) are pretrained by minimizing the cross-entropy loss for next-token prediction. In this paper, we study whether this optimization strategy can induce ge…
The Newton-Muon Optimizer
Zhehang Du, Weijie Su
The Muon optimizer has received considerable attention for its strong performance in training large language models, yet the design principle behind its matrix-gradient orthogonali…
Pretrained Multilingual Transformers Reveal Quantitative Distance Between Human Languages
Yue Zhao, Jiatao Gu, Paloma JeretiÄ +1
Understanding the distance between human languages is central to linguistics, anthropology, and tracing human evolutionary history. Yet, while linguistics has long provided rich qu…
Restoring Calibration for Aligned Large Language Models: A Calibration-Aware Fine-Tuning Approach
Jiancong Xiao, Bojian Hou, Zhanliang Wang +4
One of the key technologies for the success of Large Language Models (LLMs) is preference alignment. However, a notable side effect of preference alignment is poor calibration: whi…