Advances in Asynchronous Parallel and Distributed Optimization
arXiv:2006.13838
Abstract
Motivated by large-scale optimization problems arising in the context of machine learning, there have been several advances in the study of asynchronous parallel and distributed optimization methods during the past decade. Asynchronous methods do not require all processors to maintain a consistent view of the optimization variables. Consequently, they generally can make more efficient use of computational resources than synchronous methods, and they are not sensitive to issues like stragglers (i.e., slow nodes) and unreliable communication links. Mathematical modeling of asynchronous methods involves proper accounting of information delays, which makes their analysis challenging. This article reviews recent developments in the design and analysis of asynchronous optimization methods, covering both centralized methods, where all processors update a master copy of the optimization variables, and decentralized methods, where each processor maintains a local copy of the variables. The analysis provides insights as to how the degree of asynchrony impacts convergence rates, especially in stochastic optimization methods.
33 pages, 4 figures
References in corpus (8)
- Revisiting Distributed Synchronous SGD
- A Stochastic Gradient Method with an Exponential Convergence Rate for Finite Training Sets
- Finito: A Faster, Permutable Incremental Gradient Method for Big Data Problems
- An Asynchronous, Decentralized Solution Framework for the Large Scale Unit Commitment Problem
- Analysis and Implementation of an Asynchronous Optimization Algorithm for the Parameter Server
- Measuring scheduling efficiency of RNNs for NLP applications
- SySCD: A System-Aware Parallel Coordinate Descent Algorithm
- Variance Reduced Coordinate Descent with Acceleration: New Method With a Surprising Application to Finite-Sum Problems