5 papers
Muon is Provably Faster with Momentum Variance Reduction
Xun Qian, Hussein Rammal, Dmitry Kovalev +1
Recent empirical research has demonstrated that deep learning optimizers based on the linear minimization oracle (LMO) over specifically chosen Non-Euclidean norm balls, such as Mu…
Better LMO-based Momentum Methods with Second-Order Information
Sarit Khirirat, Abdurakhmon Sadiev, Yury Demidovich +1
The use of momentum in stochastic optimization algorithms has shown empirical success across a range of machine learning tasks. Recently, a new class of stochastic momentum algorit…
Improved Convergence in Parameter-Agnostic Error Feedback through Momentum
Abdurakhmon Sadiev, Yury Demidovich, Igor Sokolov +3
Communication compression is essential for scalable distributed training of modern machine learning models, but it often degrades convergence due to the noise it introduces. Error…
Second-order Optimization under Heavy-Tailed Noise: Hessian Clipping and Sample Complexity Limits
Abdurakhmon Sadiev, Peter Richtárik, Ilyas Fatkhullin
Heavy-tailed noise is pervasive in modern machine learning applications, arising from data heterogeneity, outliers, and non-stationary stochastic environments. While second-order m…
Bernoulli-LoRA: A Theoretical Framework for Randomized Low-Rank Adaptation
Igor Sokolov, Abdurakhmon Sadiev, Yury Demidovich +2
Parameter-efficient fine-tuning (PEFT) has emerged as a crucial approach for adapting large foundational models to specific tasks, particularly as model sizes continue to grow expo…