Trade-offs of Local SGD at Scale: An Empirical Study
arXiv:2110.08133
Abstract
As datasets and models become increasingly large, distributed training has become a necessary component to allow deep neural networks to train in reasonable amounts of time. However, distributed training can have substantial communication overhead that hinders its scalability. One strategy for reducing this overhead is to perform multiple unsynchronized SGD steps independently on each worker between synchronization steps, a technique known as local SGD. We conduct a comprehensive empirical study of local SGD and related methods on a large-scale image classification task. We find that performing local SGD comes at a price: lower communication costs (and thereby faster training) are accompanied by lower accuracy. This finding is in contrast from the smaller-scale experiments in prior work, suggesting that local SGD encounters challenges at scale. We further show that incorporating the slow momentum framework of Wang et al. (2020) consistently improves accuracy without requiring additional communication, hinting at future directions for potentially escaping this trade-off.
References in corpus (8)
- On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima
- Large Batch Training of Convolutional Networks
- Parallel training of DNNs with Natural Gradient and Parameter Averaging
- PyTorch Distributed: Experiences on Accelerating Data Parallel Training
- Parallel SGD: When does averaging help?
- Is Local SGD Better than Minibatch SGD?
- Anytime MiniBatch: Exploiting Stragglers in Online Distributed Optimization
- Selectivity considered harmful: evaluating the causal impact of class selectivity in DNNs