Showing math.OCShow all
3 papers · 1 filter
math.OC2026
Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers
Tim Tsz-Kit Lau, Weijie Su
A striking geometric disparity has long persisted in the practice of deep learning. While modern neural network architectures naturally exhibit rich symmetry and equivariance prope…
math.OC2026
Uncovering Symmetry Transfer in Large Language Models via Layer-Peeled Optimization
Zhehang Du, Hangfeng He, Weijie Su
Large language models (LLMs) are pretrained by minimizing the cross-entropy loss for next-token prediction. In this paper, we study whether this optimization strategy can induce ge…
math.OC2026
The Newton-Muon Optimizer
Zhehang Du, Weijie Su
The Muon optimizer has received considerable attention for its strong performance in training large language models, yet the design principle behind its matrix-gradient orthogonali…