11 papers
Rescaled Asynchronous SGD: Optimal Distributed Optimization under Data and System Heterogeneity
Ammar Mahran, Artavazd Maranjyan, Peter Richtárik
Asynchronous stochastic gradient descent (ASGD) is a standard way to exploit heterogeneous compute resources in distributed learning: instead of forcing fast workers to wait for sl…
Rennala MVR: Improved Time Complexity for Parallel Stochastic Optimization via Momentum-Based Variance Reduction
Zhirayr Tovmasyan, Artavazd Maranjyan, Peter Richtárik
Large-scale machine learning models are trained on clusters of machines that exhibit heterogeneous performance due to hardware variability, network delays, and system-level instabi…
Ringleader ASGD: The First Asynchronous SGD with Optimal Time Complexity under Data Heterogeneity
Artavazd Maranjyan, Peter Richtárik
Asynchronous stochastic gradient methods are central to scalable distributed optimization, particularly when devices differ in computational capabilities. Such settings arise natur…
BiCoLoR: Communication-Efficient Optimization with Bidirectional Compression and Local Training
Laurent Condat, Artavazd Maranjyan, Peter Richtárik
Slow and costly communication is often the main bottleneck in distributed optimization, especially in federated learning where it occurs over wireless networks. We introduce BiCoLo…
First Provably Optimal Asynchronous SGD for Homogeneous and Heterogeneous Data
Artavazd Maranjyan
Artificial intelligence has advanced rapidly through large neural networks trained on massive datasets using thousands of GPUs or TPUs. Such training can occupy entire data centers…
MindFlayer SGD: Efficient Parallel SGD in the Presence of Heterogeneous and Random Worker Compute Times
Artavazd Maranjyan, Omar Shaikh Omar, Peter Richtárik
We investigate the problem of minimizing the expectation of smooth nonconvex functions in a distributed setting with multiple parallel workers that are able to compute stochastic g…