4 papers
Muse: Representation Geometry of Muon Beyond Normalized Momentum
Da Chang, Qiankun Shi, Lvgang Zhang +4
Muon-style optimizers apply a polar map to matrix momentum, but their updates also depend on the representation of each parameter block before orthogonalization. We study this repr…
A Note on Stability for Orthogonalized Matrix Momentum with Client Sampling
Da Chang, Qiankun Shi, Lvgang Zhang +2
We study finite-sample generalization for a client-sampled distributed optimization scheme with matrix-valued parameters and orthogonalized momentum updates. The central quantity i…
MuonEq: Balancing Before Orthogonalization with Lightweight Equilibration
Da Chang, Qiankun Shi, Lvgang Zhang +5
Orthogonalized-update optimizers such as Muon improve training of matrix-valued parameters, but existing extensions typically either rescale updates after orthogonalization or use…
A Single-Loop Bilevel Deep Learning Method for Optimal Control of Obstacle Problems
Yongcun Song, Shangzhi Zeng, Jin Zhang +1
Optimal control of obstacle problems arises in a wide range of applications and is computationally challenging due to its nonsmoothness, nonlinearity, and bilevel structure. Classi…