large language model training 1matrix representations 1momentum methods 1optimizer geometry 1stochastic convergence 1
From the 1 of 2 linked papers with an AI index.
2 papers
cs.LG2026
Muse: Representation Geometry of Muon Beyond Normalized Momentum
Da Chang, Qiankun Shi, Lvgang Zhang +4
The paper investigates how the choice of matrix representation influences Muon-style optimizers, proposes the Muse family of optimizers that keep the same momentum and Newton–Schul…
cs.CL2025
AlphaDecay: Module-wise Weight Decay for Heavy-Tailed Balancing in LLMs
Di He, Songjun Tu, Ajay Jaiswal +4
Weight decay is a standard regularization technique for training large language models (LLMs). While it is common to assign a uniform decay rate to every layer, this approach overl…