ResIST: Layer-Wise Decomposition of ResNets for Distributed Training
arXiv:2107.00961
Abstract
We propose ResIST, a novel distributed training protocol for Residual Networks (ResNets). ResIST randomly decomposes a global ResNet into several shallow sub-ResNets that are trained independently in a distributed manner for several local iterations, before having their updates synchronized and aggregated into the global model. In the next round, new sub-ResNets are randomly generated and the process repeats until convergence. By construction, per iteration, ResIST communicates only a small portion of network parameters to each machine and never uses the full model during training. Thus, ResIST reduces the per-iteration communication, memory, and time requirements of ResNet training to only a fraction of the requirements of full-model training. In comparison to common protocols, like data-parallel training and data-parallel training with local SGD, ResIST yields a decrease in communication and compute requirements, while being competitive with respect to model performance.
26 pages, 8 figures, pre-print under review
References in corpus (13)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- Pruning Filters for Efficient ConvNets
- Large Batch Training of Convolutional Networks
- YOLO9000: Better, Faster, Stronger
- Do ImageNet Classifiers Generalize to ImageNet?
- Decentralized Stochastic Optimization and Gossip Algorithms with Compressed Communication
- Understanding and Improving Transformer From a Multi-Particle Dynamic System Point of View
- Parallel SGD: When does averaging help?
- Advances in Asynchronous Parallel and Distributed Optimization
- Layered SGD: A Decentralized and Synchronous SGD Algorithm for Scalable Deep Neural Network Training
- LAMP: Large Deep Nets with Automated Model Parallelism for Image Segmentation
- Automatic Model Parallelism for Deep Neural Networks with Compiler and Hardware Support