Local AdaAlter: Communication-Efficient Stochastic Gradient Descent with Adaptive Learning Rates
arXiv:1911.09030
Abstract
When scaling distributed training, the communication overhead is often the bottleneck. In this paper, we propose a novel SGD variant with reduced communication and adaptive learning rates. We prove the convergence of the proposed algorithm for smooth but non-convex problems. Empirical results show that the proposed algorithm significantly reduces the communication overhead, which, in turn, reduces the training time by up to 30% for the 1B word dataset.
References in corpus (21)
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- ADADELTA: An Adaptive Learning Rate Method
- Communication-Efficient Learning of Deep Networks from Decentralized Data
- Federated Learning: Strategies for Improving Communication Efficiency
- Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
- MXNet: A Flexible and Efficient Machine Learning Library for Heterogeneous Distributed Systems
- HOGWILD!: A Lock-Free Approach to Parallelizing Stochastic Gradient Descent
- QSGD: Communication-Efficient SGD via Gradient Quantization and Encoding
- Exploring the Limits of Language Modeling
- TernGrad: Ternary Gradients to Reduce Communication in Distributed Deep Learning
- Horovod: fast and easy distributed deep learning in TensorFlow
- Don't Use Large Mini-Batches, Use Local SGD
- Local SGD Converges Fast and Communicates Little
- Slow Learners are Fast
- AdaGrad stepsizes: Sharp convergence over nonconvex landscapes
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutes
- Cooperative SGD: A unified Framework for the Design and Analysis of Communication-Efficient SGD Algorithms
- On the Linear Speedup Analysis of Communication Efficient Momentum SGD for Distributed Non-Convex Optimization
- Adaptive Bound Optimization for Online Convex Optimization
- Communication-Efficient Distributed Blockwise Momentum SGD with Error-Feedback
- ImageNet Training in Minutes
Cited by in corpus (11)
- Tackling the Objective Inconsistency Problem in Heterogeneous Federated Optimization
- Adaptive Federated Optimization
- Towards Efficient and Stable K-Asynchronous Federated Learning with Unbounded Stale Gradients on Non-IID Data
- Tighter Theory for Local SGD on Identical and Heterogeneous Data
- Descending through a Crowded Valley - Benchmarking Deep Learning Optimizers
- FedCM: Federated Learning with Client-level Momentum
- Local Adaptivity in Federated Learning: Convergence and Consistency
- On the Outsized Importance of Learning Rates in Local Update Methods
- Towards Practical Adam: Non-Convexity, Convergence Theory, and Mini-Batch Acceleration
- Local AdaGrad-Type Algorithm for Stochastic Convex-Concave Optimization
- Toward Communication Efficient Adaptive Gradient Method