AutoDDL: Automatic Distributed Deep Learning with Near-Optimal Bandwidth Cost
arXiv:2301.06813 · doi:10.1109/TPDS.2024.3397800
Abstract
Recent advances in deep learning are driven by the growing scale of computation, data, and models. However, efficiently training large-scale models on distributed systems requires an intricate combination of data, operator, and pipeline parallelism, which exerts heavy burden on machine learning practitioners. To this end, we propose AutoDDL, a distributed training framework that automatically explores and exploits new parallelization schemes with near-optimal bandwidth cost. AutoDDL facilitates the description and implementation of different schemes by utilizing OneFlow's Split, Broadcast, and Partial Sum (SBP) abstraction. AutoDDL is equipped with an analytical performance model combined with a customized Coordinate Descent algorithm, which significantly reduces the scheme searching overhead. We conduct evaluations on Multi-Node-Single-GPU and Multi-Node-Multi-GPU machines using different models, including VGG and Transformer. Compared to the expert-optimized implementations, AutoDDL reduces the end-to-end training time by up to 31.1% and 10% for Transformer and up to 17.7% and 71.5% for VGG on the two parallel systems, respectively.
Published in IEEE Transactions on Parallel and Distributed Systems (TPDS), Volume: 35, Issue: 8, August 2024, doi: https://doi.org/10.1109/TPDS.2024.3397800
References in corpus (16)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- Scaling Laws for Neural Language Models
- One weird trick for parallelizing convolutional neural networks
- Training Deep Nets with Sublinear Memory Cost
- Deep Learning Recommendation Model for Personalization and Recommendation Systems
- Supporting Very Large Models using Automatic Dataflow Graph Partitioning
- Chimera: Efficiently Training Large-Scale Neural Networks with Bidirectional Pipelines
- Taming Unbalanced Training Workloads in Deep Learning with Partial Collective Operations
- OneFlow: Redesign the Distributed Deep Learning Framework from Scratch
- TeraPipe: Token-Level Pipeline Parallelism for Training Large-Scale Language Models
- Tesseract: Parallelize the Tensor Parallelism Efficiently
- Colossal-AI: A Unified Deep Learning System For Large-Scale Parallel Training
- Maximizing Parallelism in Distributed Training for Huge Neural Networks
- Colossal-Auto: Unified Automation of Parallelization and Activation Checkpoint for Large-scale Models
- Communication Bounds for Convolutional Neural Networks