GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism
arXiv:1811.06965
Abstract
Scaling up deep neural network capacity has been known as an effective approach to improving model quality for several different machine learning tasks. In many cases, increasing model capacity beyond the memory limit of a single accelerator has required developing special algorithms or infrastructure. These solutions are often architecture-specific and do not transfer to other tasks. To address the need for efficient and task-independent model parallelism, we introduce GPipe, a pipeline parallelism library that allows scaling any network that can be expressed as a sequence of layers. By pipelining different sub-sequences of layers on separate accelerators, GPipe provides the flexibility of scaling a variety of different networks to gigantic sizes efficiently. Moreover, GPipe utilizes a novel batch-splitting pipelining algorithm, resulting in almost linear speedup when a model is partitioned across multiple accelerators. We demonstrate the advantages of GPipe by training large-scale neural networks on two different tasks with distinct network architectures: (i) Image Classification: We train a 557-million-parameter AmoebaNet model and attain a top-1 accuracy of 84.4% on ImageNet-2012, (ii) Multilingual Neural Machine Translation: We train a single 6-billion-parameter, 128-layer Transformer model on a corpus spanning over 100 languages and achieve better quality than all bilingual models.
11 pages. Work in progress. Copyright 2018 by the authors
References in corpus (14)
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- Improved Regularization of Convolutional Neural Networks with Cutout
- Convolutional Sequence to Sequence Learning
- MXNet: A Flexible and Efficient Machine Learning Library for Heterogeneous Distributed Systems
- One weird trick for parallelizing convolutional neural networks
- AutoAugment: Learning Augmentation Policies from Data
- Don't Decay the Learning Rate, Increase the Batch Size
- Training Deep Nets with Sublinear Memory Cost
- Object-Part Attention Model for Fine-grained Image Classification
- Massively Multilingual Neural Machine Translation in the Wild: Findings and Challenges
- Device Placement Optimization with Reinforcement Learning
- Lingvo: a Modular and Scalable Framework for Sequence-to-Sequence Modeling
- Fixup Initialization: Residual Learning Without Normalization
- PipeDream: Fast and Efficient Pipeline Parallel DNN Training
Cited by in corpus (21)
- Towards an Effective and Efficient Deep Learning Model for COVID-19 Patterns Detection in X-ray Images
- NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM Inferencing
- The Evolution of Distributed Systems for Graph Neural Networks and their Origin in Graph Processing and Deep Learning: A Survey
- ZeRO-Infinity: Breaking the GPU Memory Wall for Extreme Scale Deep Learning
- Scalable Deep Learning on Distributed Infrastructures: Challenges, Techniques and Tools
- GSPMD: General and Scalable Parallelization for ML Computation Graphs
- On-device Training: A First Overview on Existing Systems
- High Fidelity Video Prediction with Large Stochastic Recurrent Neural Networks
- Cost-effective On-device Continual Learning over Memory Hierarchy with Miro
- Automatic Cross-Replica Sharding of Weight Update in Data-Parallel Training
- Reducing Energy Bloat in Large Model Training
- FSMoE: A Flexible and Scalable Training System for Sparse Mixture-of-Experts Models
- An Oracle for Guiding Large-Scale Model/Hybrid Parallel Training of Convolutional Neural Networks
- Deep Optimizer States: Towards Scalable Training of Transformer Models Using Interleaved Offloading
- Accelerating Neural Network Training with Distributed Asynchronous and Selective Optimization (DASO)
- LAMBERT: Layout-Aware (Language) Modeling for information extraction
- Activations and Gradients Compression for Model-Parallel Training
- Tearing Down the Memory Wall
- HyPar-Flow: Exploiting MPI and Keras for Scalable Hybrid-Parallel DNN Training using TensorFlow
- Feed-Forward Optimization With Delayed Feedback for Neural Network Training
- Trust-Aware Routing for Distributed Generative AI Inference at the Edge