Stochastic Gradient Push for Distributed Deep Learning
arXiv:1811.10792
Abstract
Distributed data-parallel algorithms aim to accelerate the training of deep neural networks by parallelizing the computation of large mini-batch gradient updates across multiple nodes. Approaches that synchronize nodes using exact distributed averaging (e.g., via AllReduce) are sensitive to stragglers and communication delays. The PushSum gossip algorithm is robust to these issues, but only performs approximate distributed averaging. This paper studies Stochastic Gradient Push (SGP), which combines PushSum with stochastic gradient updates. We prove that SGP converges to a stationary point of smooth, non-convex objectives at the same sub-linear rate as SGD, and that all nodes achieve consensus. We empirically validate the performance of SGP on image classification (ResNet-50, ImageNet) and machine translation (Transformer, WMT'16 En-De) workloads. Our code will be made publicly available.
ICML 2019
Cited by in corpus (64)
- A Unified Theory of Decentralized SGD with Changing Topology and Local Updates
- Variance-Reduced Decentralized Stochastic Optimization with Accelerated Convergence
- An improved convergence analysis for decentralized online stochastic non-convex optimization
- Communication-Efficient Distributed Deep Learning: A Comprehensive Survey
- SlowMo: Improving Communication-Efficient Distributed SGD with Slow Momentum
- A Unified and Refined Convergence Analysis for Non-Convex Decentralized Learning
- Throughput-Optimal Topology Design for Cross-Silo Federated Learning
- A Sharp Estimate on the Transient Time of Distributed Stochastic Gradient Descent
- Moniqua: Modulo Quantized Communication in Decentralized SGD
- Hyper-Sphere Quantization: Communication-Efficient SGD for Federated Learning
- Quasi-Global Momentum: Accelerating Decentralized Deep Learning on Heterogeneous Data
- Decentralized Deep Learning with Arbitrary Communication Compression
- Communication-Efficient Local Decentralized SGD Methods
- Decentralized federated learning of deep neural networks on non-iid data
- Decentralized Stochastic Gradient Tracking for Non-convex Empirical Risk Minimization
- Consensus Control for Decentralized Deep Learning
- Gossip-based Actor-Learner Architectures for Deep Reinforcement Learning
- Fully Asynchronous Distributed Optimization with Linear Convergence in Directed Networks
- Brainstorming Generative Adversarial Networks (BGANs): Towards Multi-Agent Generative Models with Distributed Private Datasets
- SPARQ-SGD: Event-Triggered and Compressed Communication in Decentralized Stochastic Optimization
- BlueFog: Make Decentralized Algorithms Practical for Optimization and Deep Learning
- Improving the Sample and Communication Complexity for Decentralized Non-Convex Optimization: A Joint Gradient Estimation and Tracking Approach
- Removing Data Heterogeneity Influence Enhances Network Topology Dependence of Decentralized SGD
- An introduction to decentralized stochastic optimization with gradient tracking
- Fast decentralized non-convex finite-sum optimization with recursive variance reduction
- Periodic Stochastic Gradient Descent with Momentum for Decentralized Training
- Advances in Asynchronous Parallel and Distributed Optimization
- PowerGossip: Practical Low-Rank Communication Compression in Decentralized Deep Learning
- Distributed Deep Learning with Event-Triggered Communication
- A Hybrid Variance-Reduced Method for Decentralized Stochastic Non-Convex Optimization
- RelaySum for Decentralized Deep Learning on Heterogeneous Data
- Gradient tracking and variance reduction for decentralized optimization and machine learning
- At Stability's Edge: How to Adjust Hyperparameters to Preserve Minima Selection in Asynchronous Training of Neural Networks?
- Federated Block Coordinate Descent Scheme for Learning Global and Personalized Models
- A Stochastic Proximal Gradient Framework for Decentralized Non-Convex Composite Optimization: Topology-Independent Sample Complexity and Communication Efficiency
- Asymptotic Network Independence in Distributed Stochastic Optimization for Machine Learning
- On the Benefits of Multiple Gossip Steps in Communication-Constrained Decentralized Optimization
- On the Convergence of Decentralized Adaptive Gradient Methods
- A Closer Look at Codistillation for Distributed Training
- D-SPIDER-SFO: A Decentralized Optimization Algorithm with Faster Convergence Rate for Nonconvex Problems
- Randomized Iterative Methods for Linear Systems: Momentum, Inexactness and Gossip
- MixML: A Unified Analysis of Weakly Consistent Parallel Learning
- Decentralized Riemannian Gradient Descent on the Stiefel Manifold
- CPR: Understanding and Improving Failure Tolerant Training for Deep Learning Recommendation with Partial Recovery
- Distributed Deep Learning in Open Collaborations
- A fast randomized incremental gradient method for decentralized non-convex optimization
- Cross-Gradient Aggregation for Decentralized Learning from Non-IID data
- Moshpit SGD: Communication-Efficient Decentralized Training on Heterogeneous Unreliable Devices
- Accelerating Gossip SGD with Periodic Global Averaging
- Overlap Local-SGD: An Algorithmic Approach to Hide Communication Delays in Distributed SGD
- Distributed Machine Learning for Wireless Communication Networks: Techniques, Architectures, and Applications
- Optimal Complexity in Decentralized Training
- Crossover-SGD: A gossip-based communication in distributed deep learning for alleviating large mini-batch problem and enhancing scalability
- Boosting Asynchronous Decentralized Learning with Model Fragmentation
- Stochastic Gradient Descent-Ascent and Consensus Optimization for Smooth Games: Convergence Analysis under Expected Co-coercivity
- A general framework for decentralized optimization with first-order methods
- Dynamic Average Diffusion with randomized Coordinate Updates
- CD-SGD: Distributed Stochastic Gradient Descent with Compression and Delay Compensation
- Improving Efficiency in Large-Scale Decentralized Distributed Training
- Distributed Stochastic Optimization With Unbounded Subgradients Over Randomly Time-Varying Networks
- Asymptotic Network Independence and Step-Size for A Distributed Subgradient Method
- Asymptotic Properties of - Method with Diminishing Stepsize
- Gradient-push algorithm for distributed optimization with event-triggered communications
- Trade-offs of Local SGD at Scale: An Empirical Study