5 papers
SILAGE: Memory-Efficient, Full-Gradient-Free Nonconvex Optimization for Nested Finite Sums
Igor Sokolov, Laurent Condat, Peter Richtárik
Empirical risk minimization on massive datasets naturally exhibits a nested double finite-sum structure, where total samples are logically or physically partitioned into …
Improved Convergence in Parameter-Agnostic Error Feedback through Momentum
Abdurakhmon Sadiev, Yury Demidovich, Igor Sokolov +3
Communication compression is essential for scalable distributed training of modern machine learning models, but it often degrades convergence due to the noise it introduces. Error…
Bernoulli-LoRA: A Theoretical Framework for Randomized Low-Rank Adaptation
Igor Sokolov, Abdurakhmon Sadiev, Yury Demidovich +2
Parameter-efficient fine-tuning (PEFT) has emerged as a crucial approach for adapting large foundational models to specific tasks, particularly as model sizes continue to grow expo…
EF21 with Bells & Whistles: Six Algorithmic Extensions of Modern Error Feedback
Ilyas Fatkhullin, Igor Sokolov, Eduard Gorbunov +2
First proposed by Seide (2014) as a heuristic, error feedback (EF) is a very popular mechanism for enforcing convergence of distributed gradient-based optimization methods enhanced…
MARINA-P: Superior Performance in Non-smooth Federated Optimization with Adaptive Stepsizes
Igor Sokolov, Peter Richtárik
Non-smooth communication-efficient federated optimization is crucial for many machine learning applications, yet remains largely unexplored theoretically. Recent advancements have…