20 papers
Convergence Analysis of Muon-type Methods with Inexact LMO in the Degenerate Case
Xun Qian, Peter Richtárik
Muon-type methods have demonstrated potentially superior performance over Adam and its variants, and have shown hyperparameter transferability across model sizes when specific norm…
SILAGE: Memory-Efficient, Full-Gradient-Free Nonconvex Optimization for Nested Finite Sums
Igor Sokolov, Laurent Condat, Peter Richtárik
Empirical risk minimization on massive datasets naturally exhibits a nested double finite-sum structure, where total samples are logically or physically partitioned into …
A Unified Primal-Dual Recipe for Accelerating Three-Operator Splitting Methods
Abdurakhmon Sadiev, Laurent Condat, Peter Richtárik
Composite optimization problems, formulated as the minimization of three functions, are ubiquitous in large-scale machine learning and signal processing. While state-of-the-art spl…
LOSCAR-SGD: Local SGD with Communication-Computation Overlap and Delay-Corrected Sparse Model Averaging
Yassine Maziane, Ammar Mahran, Artavazd Maranjyan +1
Communication is a major bottleneck in distributed learning, especially in large-scale settings and in federated learning environments with slow links. Three standard ways to reduc…
Distance-Aware Muon: Adaptive Step Scaling for Normalized Optimization
Yury Demidovich, Abhishek Chakraborty, Grigory Malinovsky +2
Muon and related normalized optimizers decouple the choice of update direction from the choice of step scale, but their practical performance remains sensitive to the scale of the…
Ringmaster LMO: Asynchronous Linear Minimization Oracle Momentum Method
Abdurakhmon Sadiev, Artavazd Maranjyan, Ivan Ilin +1
Muon has recently emerged as a strong alternative to AdamW for training neural networks, with encouraging large-scale pretraining results and growing evidence that matrix-structure…