8 papers
Convergence Analysis of Muon-type Methods with Inexact LMO in the Degenerate Case
Xun Qian, Peter Richtárik
Muon-type methods have demonstrated potentially superior performance over Adam and its variants, and have shown hyperparameter transferability across model sizes when specific norm…
Broximal Alignment for Global Non-Convex Optimization
Kaja Gruntkowska, Hanmin Li, Xun Qian +1
Most non-convex optimization theory is built around gradient dynamics, leaving global convergence largely unexplored. The dominant paradigm focuses on stationarity, certifying only…
Communication-Efficient Gluon in Federated Learning
Xun Qian, Alexander Gaponov, Grigory Malinovsky +1
Recent developments have shown that Muon-type optimizers based on linear minimization oracles (LMOs) over non-Euclidean norm balls have the potential to get superior practical perf…
A Stochastic Block-coordinate Proximal Newton Method for Nonconvex Composite Minimization
Hong Zhu, Xun Qian
This paper presents a stochastic block-coordinate proximal Newton method for minimizing the sum of a blockwise Lipschitz-continuously differentiable function and a separable nonsmo…
Muon is Provably Faster with Momentum Variance Reduction
Xun Qian, Hussein Rammal, Dmitry Kovalev +1
Recent empirical research has demonstrated that deep learning optimizers based on the linear minimization oracle (LMO) over specifically chosen Non-Euclidean norm balls, such as Mu…
A matrix-free interior point continuous trajectory for linearly constrained convex programming
Xun Qian, Li-Zhi Liao, Jie Sun
Interior point methods for solving linearly constrained convex programming involve a variable projection matrix at each iteration to deal with the linear constraints. This matrix o…