10 papers
A Unified Primal-Dual Recipe for Accelerating Three-Operator Splitting Methods
Abdurakhmon Sadiev, Laurent Condat, Peter Richtárik
Composite optimization problems, formulated as the minimization of three functions, are ubiquitous in large-scale machine learning and signal processing. While state-of-the-art spl…
Ringmaster LMO: Asynchronous Linear Minimization Oracle Momentum Method
Abdurakhmon Sadiev, Artavazd Maranjyan, Ivan Ilin +1
Muon has recently emerged as a strong alternative to AdamW for training neural networks, with encouraging large-scale pretraining results and growing evidence that matrix-structure…
A Nesterov-Accelerated Primal-Dual Splitting Algorithm for Convex Nonsmooth Optimization
Laurent Condat, Abdurakhmon Sadiev, Peter Richtárik
We investigate the integration of Nesterov-type acceleration into primal-dual methods for structured convex optimization. While proximal splitting algorithms efficiently handle com…
Tight Lower Bounds and Optimal Algorithms for Stochastic Nonconvex Optimization with Heavy-Tailed Noise
Adrien Fradin, Abdurakhmon Sadiev, Laurent Condat +1
We study stochastic nonconvex optimization under heavy-tailed noise. In this setting, the stochastic gradients only have bounded -th central moment (-BCM) for some $p \in (1,…
Better LMO-based Momentum Methods with Second-Order Information
Sarit Khirirat, Abdurakhmon Sadiev, Yury Demidovich +1
The use of momentum in stochastic optimization algorithms has shown empirical success across a range of machine learning tasks. Recently, a new class of stochastic momentum algorit…
Improved Convergence in Parameter-Agnostic Error Feedback through Momentum
Abdurakhmon Sadiev, Yury Demidovich, Igor Sokolov +3
Communication compression is essential for scalable distributed training of modern machine learning models, but it often degrades convergence due to the noise it introduces. Error…