7 papers
Tighter Performance Theory of FedExProx
Wojciech Anyszka, Kaja Gruntkowska, Alexander Tyurin +1
We revisit FedExProx - a recently proposed distributed optimization method designed to enhance convergence properties of parallel proximal algorithms via extrapolation. In the proc…
Local SGD and Federated Averaging Through the Lens of Time Complexity
Adrien Fradin, Peter Richtárik, Alexander Tyurin
We revisit the classical Local SGD and Federated Averaging (FedAvg) methods for distributed optimization and federated learning. While prior work has primarily focused on iteration…
Ringmaster ASGD: The First Asynchronous SGD with Optimal Time Complexity
Artavazd Maranjyan, Alexander Tyurin, Peter Richtárik
Asynchronous Stochastic Gradient Descent (Asynchronous SGD) is a cornerstone method for parallelizing learning in distributed machine learning. However, its performance suffers und…
On the Optimal Time Complexities in Decentralized Stochastic Asynchronous Optimization
Alexander Tyurin, Peter Richtárik
We consider the decentralized stochastic asynchronous optimization setup, where many workers asynchronously calculate stochastic gradients and asynchronously communicate with each…
Freya PAGE: First Optimal Time Complexity for Large-Scale Nonconvex Finite-Sum Optimization with Heterogeneous Asynchronous Computations
Alexander Tyurin, Kaja Gruntkowska, Peter Richtárik
In practical distributed systems, workers are typically not homogeneous, and due to differences in hardware configurations and network conditions, can have highly varying processin…
Improving the Worst-Case Bidirectional Communication Complexity for Nonconvex Distributed Optimization under Function Similarity
Kaja Gruntkowska, Alexander Tyurin, Peter Richtárik
Effective communication between the server and workers plays a key role in distributed optimization. In this paper, we focus on optimizing the server-to-worker communication, uncov…