4 papers
Muse: Representation Geometry of Muon Beyond Normalized Momentum
Da Chang, Qiankun Shi, Lvgang Zhang +4
Muon-style optimizers apply a polar map to matrix momentum, but their updates also depend on the representation of each parameter block before orthogonalization. We study this repr…
A Fletcher's Augmented Lagrangian-Based Stochastic First-Order Method for Nonconvex Equality-Constrained Optimization
Yawen Cui, Qiankun Shi, Xiao Wang +1
In this paper, we study nonconvex equality-constrained optimization problems in which only stochastic first-order approximations of the objective and constraint functions are avail…
Adaptive directional decomposition methods for nonconvex constrained optimization
Qiankun Shi, Xiao Wang
In this paper, we study nonconvex constrained optimization problems with both equality and inequality constraints, covering deterministic and stochastic settings. We propose a nove…
Optimal Complexity in Byzantine-Robust Distributed Stochastic Optimization with Data Heterogeneity
Qiankun Shi, Jie Peng, Kun Yuan +2
In this paper, we establish tight lower bounds for Byzantine-robust distributed first-order stochastic optimization methods in both strongly convex and non-convex stochastic optimi…