Near-Optimal Sparse Allreduce for Distributed Deep Learning
arXiv:2201.07598 · doi:10.1145/3503221.3508399
Abstract
Communication overhead is one of the major obstacles to train large deep learning models at scale. Gradient sparsification is a promising technique to reduce the communication volume. However, it is very challenging to obtain real performance improvement because of (1) the difficulty of achieving an scalable and efficient sparse allreduce algorithm and (2) the sparsification overhead. This paper proposes O-Top, a scheme for distributed training with sparse gradients. O-Top integrates a novel sparse allreduce algorithm (less than 6 communication volume which is asymptotically optimal) with the decentralized parallel Stochastic Gradient Descent (SGD) optimizer, and its convergence is proved. To reduce the sparsification overhead, O-Top efficiently selects the top- gradient values according to an estimated threshold. Evaluations are conducted on the Piz Daint supercomputer with neural network models from different deep learning domains. Empirical results show that O-Top achieves similar model accuracy to dense allreduce. Compared with the optimized dense and the state-of-the-art sparse allreduces, O-Top is more scalable and significantly improves training throughput (e.g., 3.29x-12.95x improvement for BERT on 256 GPUs).
Published in Proceedings of the 27th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming (PPoPP'22), April 2-6, 2022, Pages 135-149, https://doi.org/10.1145/3503221.3508399
References in corpus (15)
- Adam: A Method for Stochastic Optimization
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
- QSGD: Communication-Efficient SGD via Gradient Quantization and Encoding
- TernGrad: Ternary Gradients to Reduce Communication in Distributed Deep Learning
- Sparsity in Deep Learning: Pruning and growth for efficient inference and training in neural networks
- Demystifying Parallel and Distributed Deep Learning: An In-Depth Concurrency Analysis
- signSGD: Compressed Optimisation for Non-Convex Problems
- Stochastic Distributed Learning with Gradient Quantization and Variance Reduction
- Understanding Top-k Sparsification in Distributed Deep Learning
- Memory-Efficient Pipeline-Parallel DNN Training
- Taming Unbalanced Training Workloads in Deep Learning with Partial Collective Operations
- Building Efficient ConvNets using Redundant Feature Pruning
- DAPPLE: A Pipelined Data Parallel Approach for Training Large Models
- Asynchronous Decentralized SGD with Quantized and Local Updates