Chimera: Efficiently Training Large-Scale Neural Networks with Bidirectional Pipelines
arXiv:2107.06925 · doi:10.1145/3458817.3476145
Abstract
Training large deep learning models at scale is very challenging. This paper proposes Chimera, a novel pipeline parallelism scheme which combines bidirectional pipelines for efficiently training large-scale models. Chimera is a synchronous approach and therefore no loss of accuracy, which is more convergence-friendly than asynchronous approaches. Compared with the latest synchronous pipeline approach, Chimera reduces the number of bubbles by up to 50%; benefiting from the sophisticated scheduling of bidirectional pipelines, Chimera has a more balanced activation memory consumption. Evaluations are conducted on Transformer based language models. For a GPT-2 model with 1.3 billion parameters running on 2,048 GPU nodes of the Piz Daint supercomputer, Chimera improves the training throughput by 1.16x-2.34x over the state-of-the-art synchronous and asynchronous pipeline approaches.
Published in Proceedings of the 2021 International Conference for High Performance Computing, Networking, Storage and Analysis (SC'21), November 2021, Article No.: 27, Pages 1-14. Best Paper Finalist
References in corpus (16)
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- HuggingFace's Transformers: State-of-the-art Natural Language Processing
- Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
- One weird trick for parallelizing convolutional neural networks
- QSGD: Communication-Efficient SGD via Gradient Quantization and Encoding
- Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
- Training Deep Nets with Sublinear Memory Cost
- Sparsity in Deep Learning: Pruning and growth for efficient inference and training in neural networks
- Demystifying Parallel and Distributed Deep Learning: An In-Depth Concurrency Analysis
- Communication-Efficient Distributed Deep Learning: A Comprehensive Survey
- Memory-Efficient Pipeline-Parallel DNN Training
- Taming Unbalanced Training Workloads in Deep Learning with Partial Collective Operations
- DAPPLE: A Pipelined Data Parallel Approach for Training Large Models
- AMPNet: Asynchronous Model-Parallel Training for Dynamic Neural Networks
- Asynchronous Decentralized SGD with Quantized and Local Updates
- Red-blue pebbling revisited: near optimal parallel matrix-matrix multiplication
Cited by in corpus (9)
- Merak: An Efficient Distributed DNN Training Framework with Automated 3D Parallelism for Giant Foundation Models
- Oobleck: Resilient Distributed Training of Large Models Using Pipeline Templates
- Hanayo: Harnessing Wave-like Pipeline Parallelism for Enhanced Large Model Training Efficiency
- Efficient Quantized Sparse Matrix Operations on Tensor Cores
- ReCycle: Resilient Training of Large DNNs using Pipeline Adaptation
- FuncPipe: A Pipelined Serverless Framework for Fast and Cost-efficient Training of Deep Learning Models
- AutoDDL: Automatic Distributed Deep Learning with Near-Optimal Bandwidth Cost
- DIP: Efficient Large Multimodal Model Training with Dynamic Interleaved Pipeline
- HARP: Orchestrating Automated Parallel Training on Heterogeneous GPU Clusters