works on

From the 1 of 8 linked papers with an AI index.

collaborators

8 papers

cs.LG2026

Muse: Representation Geometry of Muon Beyond Normalized Momentum

Da Chang, Qiankun Shi, Lvgang Zhang +4

The paper investigates how the choice of matrix representation influences Muon-style optimizers, proposes the Muse family of optimizers that keep the same momentum and Newton–Schul…

math.OC2026

A Fletcher's Augmented Lagrangian-Based Stochastic First-Order Method for Nonconvex Equality-Constrained Optimization

Yawen Cui, Qiankun Shi, Xiao Wang +1

In this paper, we study nonconvex equality-constrained optimization problems in which only stochastic first-order approximations of the objective and constraint functions are avail…

cs.LG2026

The Effective Number of Nonzeros: Theory and Regularization for Sparse Recovery

Haoyu He, Hao Wang, Jiashan Wang +1

Classical sparse recovery treats all nonzero entries equally, though numerical noise often creates long tails of negligible coefficients. This paper develops an entropy-based notio…

cs.LG2026

A Note on Stability for Orthogonalized Matrix Momentum with Client Sampling

Da Chang, Qiankun Shi, Lvgang Zhang +2

We study finite-sample generalization for a client-sampled distributed optimization scheme with matrix-valued parameters and orthogonalized momentum updates. The central quantity i…

cs.LG2026

MuonEq: Balancing Before Orthogonalization with Lightweight Equilibration

Da Chang, Qiankun Shi, Lvgang Zhang +5

Orthogonalized-update optimizers such as Muon improve training of matrix-valued parameters, but existing extensions typically either rescale updates after orthogonalization or use…

math.OC2025

Adaptive directional decomposition methods for nonconvex constrained optimization

Qiankun Shi, Xiao Wang

In this paper, we study nonconvex constrained optimization problems with both equality and inequality constraints, covering deterministic and stochastic settings. We propose a nove…