4 papers
Rank-Efficient LoRA via Joint Tangent-Space Optimization under Isotropic Curvature
Zihan Zhu, Zhehang Du, Xuyang Chen +5
Low-Rank Adaptation (LoRA) is an effective approach for adapting large pretrained models by learning low-rank weight updates. In practice, the LoRA rank is used to control an adapt…
Agents' Last Exam
Yiyou Sun, Xinyang Han, Weichen Zhang +306
Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deployment across many professional d…
Uncovering Symmetry Transfer in Large Language Models via Layer-Peeled Optimization
Zhehang Du, Hangfeng He, Weijie Su
Large language models (LLMs) are pretrained by minimizing the cross-entropy loss for next-token prediction. In this paper, we study whether this optimization strategy can induce ge…
The Newton-Muon Optimizer
Zhehang Du, Weijie Su
The Muon optimizer has received considerable attention for its strong performance in training large language models, yet the design principle behind its matrix-gradient orthogonali…