Communication-Efficient Distributed Deep Learning: A Comprehensive Survey
arXiv:2003.06307
Abstract
Distributed deep learning (DL) has become prevalent in recent years to reduce training time by leveraging multiple computing devices (e.g., GPUs/TPUs) due to larger models and datasets. However, system scalability is limited by communication becoming the performance bottleneck. Addressing this communication issue has become a prominent research topic. In this paper, we provide a comprehensive survey of the communication-efficient distributed training algorithms, focusing on both system-level and algorithmic-level optimizations. We first propose a taxonomy of data-parallel distributed training algorithms that incorporates four primary dimensions: communication synchronization, system architectures, compression techniques, and parallelism of communication and computing tasks. We then investigate state-of-the-art studies that address problems in these four dimensions. We also compare the convergence rates of different algorithms to understand their convergence speed. Additionally, we conduct extensive experiments to empirically compare the convergence performance of various mainstream distributed training algorithms. Based on our system-level communication cost analysis, theoretical and experimental convergence speed comparison, we provide readers with an understanding of which algorithms are more efficient under specific distributed environments. Our research also extrapolates potential directions for further optimizations.
References in corpus (19)
- Quantized Neural Networks: Training Neural Networks with Low Precision Weights and Activations
- cuDNN: Efficient Primitives for Deep Learning
- One weird trick for parallelizing convolutional neural networks
- Revisiting Distributed Synchronous SGD
- Device Placement Optimization with Reinforcement Learning
- Decentralized Stochastic Optimization and Gossip Algorithms with Compressed Communication
- Error Feedback Fixes SignSGD and other Gradient Compression Schemes
- On the Linear Speedup Analysis of Communication Efficient Momentum SGD for Distributed Non-Convex Optimization
- A Survey of Deep Learning Techniques for Neural Machine Translation
- AdaComp : Adaptive Residual Gradient Compression for Data-Parallel Distributed Training
- Understanding Top-k Sparsification in Distributed Deep Learning
- How to scale distributed deep learning?
- Priority-based Parameter Propagation for Distributed DNN Training
- PASSCoDe: Parallel ASynchronous Stochastic dual Co-ordinate Descent
- On the Computation and Communication Complexity of Parallel SGD with Dynamic Batch Sizes for Stochastic Non-Convex Optimization
- Blink: Fast and Generic Collectives for Distributed ML
- Communication-Efficient Decentralized Learning with Sparsification and Adaptive Peer Selection
- Communication Contention Aware Scheduling of Multiple Deep Learning Training Jobs
- Efficient Use of Limited-Memory Accelerators for Linear Learning on Heterogeneous Systems
Cited by in corpus (15)
- Chimera: Efficiently Training Large-Scale Neural Networks with Bidirectional Pipelines
- Privacy-preserving Artificial Intelligence Techniques in Biomedicine
- An Overview of Federated Learning at the Edge and Distributed Ledger Technologies for Robotic and Autonomous Systems
- The Evolution of Distributed Systems for Graph Neural Networks and their Origin in Graph Processing and Deep Learning: A Survey
- Towards Efficient Synchronous Federated Training: A Survey on System Optimization Strategies
- On the Utility of Gradient Compression in Distributed Training Systems
- Local SGD With a Communication Overhead Depending Only on the Number of Workers
- Straggler-Resilient Distributed Machine Learning with Dynamic Backup Workers
- Large-Scale Deep Learning Optimizations: A Comprehensive Survey
- Communication-efficient SGD: From Local SGD to One-Shot Averaging
- More Industry-friendly: Federated Learning with High Efficient Design
- ResIST: Layer-Wise Decomposition of ResNets for Distributed Training
- Towards a Standardized Representation for Deep Learning Collective Algorithms
- Communication-Efficient Distributed Learning via Sparse and Adaptive Stochastic Gradient
- Privacy-Preserving Serverless Edge Learning with Decentralized Small Data