Parallax: Sparsity-aware Data Parallel Training of Deep Neural Networks
arXiv:1808.02621
Abstract
The employment of high-performance servers and GPU accelerators for training deep neural network models have greatly accelerated recent advances in deep learning (DL). DL frameworks, such as TensorFlow, MXNet, and Caffe2, have emerged to assist DL researchers to train their models in a distributed manner. Although current DL frameworks scale well for image classification models, there remain opportunities for scalable distributed training on natural language processing (NLP) models. We found that current frameworks show relatively low scalability on training NLP models due to the lack of consideration to the difference in sparsity of model parameters. In this paper, we propose Parallax, a framework that optimizes data parallel training by utilizing the sparsity of model parameters. Parallax introduces a hybrid approach that combines Parameter Server and AllReduce architectures to optimize the amount of data transfer according to the sparsity. Experiments show that Parallax built atop TensorFlow achieves scalable training throughput on both dense and sparse models while requiring little effort from its users. Parallax achieves up to 2.8x, 6.02x speedup for NLP models than TensorFlow and Horovod with 48 GPUs, respectively. The training speed for the image classification models is equal to Horovod and 1.53x faster than TensorFlow.
13 pages, 9 figures
References in corpus (18)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Semi-Supervised Classification with Graph Convolutional Networks
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
- Conditional Image Generation with PixelCNN Decoders
- MXNet: A Flexible and Efficient Machine Learning Library for Heterogeneous Distributed Systems
- QSGD: Communication-Efficient SGD via Gradient Quantization and Encoding
- Exploring the Limits of Language Modeling
- Revisiting Distributed Synchronous SGD
- End-to-end Continuous Speech Recognition using Attention-based Recurrent NN: First Results
- Highly Scalable Deep Learning Training System with Mixed-Precision: Training ImageNet in Four Minutes
- Extremely Large Minibatch SGD: Training ResNet-50 on ImageNet in 15 Minutes
- Supporting Very Large Models using Automatic Dataflow Graph Partitioning
- PipeDream: Fast and Efficient Pipeline Parallel DNN Training
- AutoPruner: An End-to-End Trainable Filter Pruning Method for Efficient Deep Model Inference
- MXNET-MPI: Embedding MPI parallelism in Parameter Server Task Model for scaling Deep Learning
- Unifying Data, Model and Hybrid Parallelism in Deep Learning via Tensor Tiling
- Efficient Training of Convolutional Neural Nets on Large Distributed Systems