3 papers
cs.LG2026
Muse: Representation Geometry of Muon Beyond Normalized Momentum
Da Chang, Qiankun Shi, Lvgang Zhang +4
Muon-style optimizers apply a polar map to matrix momentum, but their updates also depend on the representation of each parameter block before orthogonalization. We study this repr…
cs.CL2025
AlphaDecay: Module-wise Weight Decay for Heavy-Tailed Balancing in LLMs
Di He, Songjun Tu, Ajay Jaiswal +4
Weight decay is a standard regularization technique for training large language models (LLMs). While it is common to assign a uniform decay rate to every layer, this approach overl…
cs.DS2024
Block Coordinate Descent Methods for Optimization under J-Orthogonality Constraints with Applications
Di He, Ganzhao Yuan, Xiao Wang +1
The J-orthogonal matrix, also referred to as the hyperbolic orthogonal matrix, is a class of special orthogonal matrix in hyperbolic space, notable for its advantageous properties.…