collaborators

5 papers

math.OC2025

Muon is Provably Faster with Momentum Variance Reduction

Xun Qian, Hussein Rammal, Dmitry Kovalev +1

Recent empirical research has demonstrated that deep learning optimizers based on the linear minimization oracle (LMO) over specifically chosen Non-Euclidean norm balls, such as Mu…

math.OC2025

Better LMO-based Momentum Methods with Second-Order Information

Sarit Khirirat, Abdurakhmon Sadiev, Yury Demidovich +1

The use of momentum in stochastic optimization algorithms has shown empirical success across a range of machine learning tasks. Recently, a new class of stochastic momentum algorit…

math.OC2025

Improved Convergence in Parameter-Agnostic Error Feedback through Momentum

Abdurakhmon Sadiev, Yury Demidovich, Igor Sokolov +3

Communication compression is essential for scalable distributed training of modern machine learning models, but it often degrades convergence due to the noise it introduces. Error…

math.OC2025

Second-order Optimization under Heavy-Tailed Noise: Hessian Clipping and Sample Complexity Limits

Abdurakhmon Sadiev, Peter Richtárik, Ilyas Fatkhullin

Heavy-tailed noise is pervasive in modern machine learning applications, arising from data heterogeneity, outliers, and non-stationary stochastic environments. While second-order m…

cs.LG2025

Bernoulli-LoRA: A Theoretical Framework for Randomized Low-Rank Adaptation

Igor Sokolov, Abdurakhmon Sadiev, Yury Demidovich +2

Parameter-efficient fine-tuning (PEFT) has emerged as a crucial approach for adapting large foundational models to specific tasks, particularly as model sizes continue to grow expo…