6 papers · 1 filter
Rank-Efficient LoRA via Joint Tangent-Space Optimization under Isotropic Curvature
Zihan Zhu, Zhehang Du, Xuyang Chen +5
Low-Rank Adaptation (LoRA) is an effective approach for adapting large pretrained models by learning low-rank weight updates. In practice, the LoRA rank is used to control an adapt…
Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers
Tim Tsz-Kit Lau, Weijie Su
A striking geometric disparity has long persisted in the practice of deep learning. While modern neural network architectures naturally exhibit rich symmetry and equivariance prope…
When Does Model Collapse Occur in Structured Interactive Learning?
Yuchen Wu, Kangjie Zhou, Weijie Su
The proliferation of generative artificial intelligence has given rise to an interactive learning environment, where model parameters are continuously updated using not only data g…
Uncovering Symmetry Transfer in Large Language Models via Layer-Peeled Optimization
Zhehang Du, Hangfeng He, Weijie Su
Large language models (LLMs) are pretrained by minimizing the cross-entropy loss for next-token prediction. In this paper, we study whether this optimization strategy can induce ge…
The Newton-Muon Optimizer
Zhehang Du, Weijie Su
The Muon optimizer has received considerable attention for its strong performance in training large language models, yet the design principle behind its matrix-gradient orthogonali…
Pretrained Multilingual Transformers Reveal Quantitative Distance Between Human Languages
Yue Zhao, Jiatao Gu, Paloma Jeretič +1
Understanding the distance between human languages is central to linguistics, anthropology, and tracing human evolutionary history. Yet, while linguistics has long provided rich qu…