Asynchronous Decentralized Parallel Stochastic Gradient Descent
arXiv:1710.06952
Abstract
Most commonly used distributed machine learning systems are either synchronous or centralized asynchronous. Synchronous algorithms like AllReduce-SGD perform poorly in a heterogeneous environment, while asynchronous algorithms using a parameter server suffer from 1) communication bottleneck at parameter servers when workers are many, and 2) significantly worse convergence when the traffic to parameter server is congested. Can we design an algorithm that is robust in a heterogeneous environment, while being communication efficient and maintaining the best-possible convergence rate? In this paper, we propose an asynchronous decentralized stochastic gradient decent algorithm (AD-PSGD) satisfying all above expectations. Our theoretical analysis shows AD-PSGD converges at the optimal rate as SGD and has linear speedup w.r.t. number of workers. Empirically, AD-PSGD outperforms the best of decentralized parallel SGD (D-PSGD), asynchronous parallel SGD (A-PSGD), and standard data parallel SGD (AllReduce-SGD), often by orders of magnitude in a heterogeneous environment. When training ResNet-50 on ImageNet with up to 128 GPUs, AD-PSGD converges (w.r.t epochs) similarly to the AllReduce-SGD, but each epoch can be up to 4-8X faster than its synchronous counterparts in a network-sharing HPC environment. To the best of our knowledge, AD-PSGD is the first asynchronous algorithm that achieves a similar epoch-wise convergence rate as AllReduce-SGD, at an over 100-GPU scale.
References in corpus (6)
- Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
- A decentralized proximal-gradient method with network independent step-sizes and separated convergence rates
- D: Decentralized Training over Decentralized Data
- DSA: Decentralized Double Stochastic Averaging Gradient Algorithm
- Decentralized Consensus Optimization with Asynchrony and Delays
- Decentralized RLS with Data-Adaptive Censoring for Regressions over Large-Scale Networks
Cited by in corpus (20)
- Asynchronous Federated Optimization
- D: Decentralized Training over Decentralized Data
- Adaptive Federated Learning in Resource Constrained Edge Computing Systems
- HET: Scaling out Huge Embedding Model Training via Cache-enabled Distributed Framework
- MATCHA: Speeding Up Decentralized SGD via Matching Decomposition Sampling
- Asynchronous Accelerated Proximal Stochastic Gradient for Strongly Convex Distributed Finite Sums
- PIRATE: A Blockchain-based Secure Framework of Distributed Machine Learning in 5G Networks
- Robust and Communication-Efficient Collaborative Learning
- At Stability's Edge: How to Adjust Hyperparameters to Preserve Minima Selection in Asynchronous Training of Neural Networks?
- DBS: Dynamic Batch Size For Distributed Deep Neural Network Training
- Asynchronous Decentralized Learning of a Neural Network
- Asynchrony and Acceleration in Gossip Algorithms
- Accelerated Sparsified SGD with Error Feedback
- Practical Federated Learning without a Server
- Hogwild! over Distributed Local Data Sets with Linearly Increasing Mini-Batch Sizes
- Accelerating MoE Model Inference with Expert Sharding
- Boosting Asynchronous Decentralized Learning with Model Fragmentation
- Pushing the boundaries of parallel Deep Learning -- A practical approach
- Nested Distributed Gradient Methods with Stochastic Computation Errors
- Task allocation for decentralized training in heterogeneous environment