TensorOpt: Exploring the Tradeoffs in Distributed DNN Training with Auto-Parallelism
arXiv:2004.10856 · doi:10.1109/TPDS.2021.3132413
Abstract
A good parallelization strategy can significantly improve the efficiency or reduce the cost for the distributed training of deep neural networks (DNNs). Recently, several methods have been proposed to find efficient parallelization strategies but they all optimize a single objective (e.g., execution time, memory consumption) and produce only one strategy. We propose FT, an efficient algorithm that searches for an optimal set of parallelization strategies to allow the trade-off among different objectives. FT can adapt to different scenarios by minimizing the memory consumption when the number of devices is limited and fully utilize additional resources to reduce the execution time. For popular DNN models (e.g., vision, language), an in-depth analysis is conducted to understand the trade-offs among different objectives and their influence on the parallelization strategies. We also develop a user-friendly system, called TensorOpt, which allows users to run their distributed DNN training jobs without caring the details of parallelization strategies. Experimental results show that FT runs efficiently and provides accurate estimation of runtime costs, and TensorOpt is more flexible in adapting to resource availability compared with existing frameworks.
References in corpus (9)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- One weird trick for parallelizing convolutional neural networks
- Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
- GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
- Device Placement Optimization with Reinforcement Learning
- Supporting Very Large Models using Automatic Dataflow Graph Partitioning
- Beyond Data and Model Parallelism for Deep Neural Networks
- A Graph Theoretic Framework of Recomputation Algorithms for Memory-Efficient Backpropagation
Cited by in corpus (5)
- Galvatron: Efficient Transformer Training over Multiple GPUs Using Automatic Parallelism
- Local Critic Training for Model-Parallel Learning of Deep Neural Networks
- ReCycle: Resilient Training of Large DNNs using Pipeline Adaptation
- Improving Automatic Parallel Training via Balanced Memory Workload Optimization
- MLP-Offload: Multi-Level, Multi-Path Offloading for LLM Pre-training to Break the GPU Memory Wall