11 citations · 42 across the 15 of their papers we have counts for
10 papers · 1 filter
Convergence Analysis of Muon-type Methods with Inexact LMO in the Degenerate Case
Xun Qian, Peter Richtárik
Muon-type methods have demonstrated potentially superior performance over Adam and its variants, and have shown hyperparameter transferability across model sizes when specific norm…
Broximal Alignment for Global Non-Convex Optimization
Kaja Gruntkowska, Hanmin Li, Xun Qian +1
Most non-convex optimization theory is built around gradient dynamics, leaving global convergence largely unexplored. The dominant paradigm focuses on stationarity, certifying only…
Muon is Provably Faster with Momentum Variance Reduction
Xun Qian, Hussein Rammal, Dmitry Kovalev +1
Recent empirical research has demonstrated that deep learning optimizers based on the linear minimization oracle (LMO) over specifically chosen Non-Euclidean norm balls, such as Mu…
A matrix-free interior point continuous trajectory for linearly constrained convex programming
Xun Qian, Li-Zhi Liao, Jie Sun
Interior point methods for solving linearly constrained convex programming involve a variable projection matrix at each iteration to deal with the linear constraints. This matrix o…
A Stochastic Block-coordinate Proximal Newton Method for Nonconvex Composite Minimization
Hong Zhu, Xun Qian
This paper presents a stochastic block-coordinate proximal Newton method for minimizing the sum of a blockwise Lipschitz-continuously differentiable function and a separable nonsmo…
Error Compensated Loopless SVRG, Quartz, and SDCA for Distributed Optimization
Xun Qian, Hanze Dong, Peter Richtárik +1
The communication of gradients is a key bottleneck in distributed training of large scale machine learning models. In order to reduce the communication cost, gradient compression (…