Efficient and Robust Parallel DNN Training through Model Parallelism on Multi-GPU Platform
arXiv:1809.02839
Abstract
The training process of Deep Neural Network (DNN) is compute-intensive, often taking days to weeks to train a DNN model. Therefore, parallel execution of DNN training on GPUs is a widely adopted approach to speed up the process nowadays. Due to the implementation simplicity, data parallelism is currently the most commonly used parallelization method. Nonetheless, data parallelism suffers from excessive inter-GPU communication overhead due to frequent weight synchronization among GPUs. Another approach is pipelined model parallelism, which partitions a DNN model among GPUs, and processes multiple mini-batches concurrently. This approach can significantly reduce inter-GPU communication cost compared to data parallelism. However, pipelined model parallelism faces the weight staleness issue; that is, gradients are computed with stale weights, leading to training instability and accuracy loss. In this paper, we present a pipelined model parallel execution method that enables high GPU utilization while maintaining robust training accuracy via a novel weight prediction technique, SpecTrain. Experimental results show that our proposal achieves up to 8.91x speedup compared to data parallelism on a 4-GPU platform while maintaining comparable model accuracy.
References in corpus (15)
- Sequence to Sequence Learning with Neural Networks
- Caffe: Convolutional Architecture for Fast Feature Embedding
- On the difficulty of training Recurrent Neural Networks
- DoReFa-Net: Training Low Bitwidth Convolutional Neural Networks with Low Bitwidth Gradients
- cuDNN: Efficient Primitives for Deep Learning
- Deep Gradient Compression: Reducing the Communication Bandwidth for Distributed Training
- One weird trick for parallelizing convolutional neural networks
- Revisiting Distributed Synchronous SGD
- Massively Multilingual Neural Machine Translation in the Wild: Findings and Challenges
- Device Placement Optimization with Reinforcement Learning
- Towards Understanding Generalization of Deep Learning: Perspective of Loss Landscapes
- PipeDream: Fast and Efficient Pipeline Parallel DNN Training
- Decoupled Parallel Backpropagation with Convergence Guarantee
- Training Neural Networks Using Features Replay
- MXNET-MPI: Embedding MPI parallelism in Parameter Server Task Model for scaling Deep Learning
Cited by in corpus (17)
- Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
- Click-Through Rate Prediction in Online Advertising: A Literature Review
- XPipe: Efficient Pipeline Model Parallelism for Multi-GPU DNN Training
- DAPPLE: A Pipelined Data Parallel Approach for Training Large Models
- Efficient Pipeline Planning for Expedited Distributed DNN Training
- Orchestrating the Development Lifecycle of Machine Learning-Based IoT Applications: A Taxonomy and Survey
- Taming Momentum in a Distributed Asynchronous Environment
- Pipelined Backpropagation at Scale: Training Large Models without Batches
- Amazon SageMaker Model Parallelism: A General and Flexible Framework for Large Model Training
- Parareal Neural Networks Emulating a Parallel-in-time Algorithm
- Distributed Machine Learning for Computational Engineering using MPI
- A Study of Checkpointing in Large Scale Training of Deep Neural Networks
- Automatic Graph Partitioning for Very Large-scale Deep Learning
- Thanks for Nothing: Predicting Zero-Valued Activations with Lightweight Convolutional Neural Networks
- Integrating Deep Learning in Domain Sciences at Exascale
- Adaptive Braking for Mitigating Gradient Delay
- ResIST: Layer-Wise Decomposition of ResNets for Distributed Training