Improving Automatic Parallel Training via Balanced Memory Workload Optimization
arXiv:2307.02031 · doi:10.1109/TKDE.2024.3370614
Abstract
Transformer models have emerged as the leading approach for achieving state-of-the-art performance across various application domains, serving as the foundation for advanced large-scale deep learning (DL) models. However, efficiently training these models across multiple GPUs remains a complex challenge due to the abundance of parallelism options. Existing DL systems either require manual efforts to design distributed training plans or limit parallelism combinations to a constrained search space. In this paper, we present Galvatron-BMW, a novel system framework that integrates multiple prevalent parallelism dimensions and automatically identifies the most efficient hybrid parallelism strategy. To effectively navigate this vast search space, we employ a decision tree approach for decomposition and pruning based on intuitive insights. We further utilize a dynamic programming search algorithm to derive the optimal plan. Moreover, to improve resource utilization and enhance system efficiency, we propose a bi-objective optimization workflow that focuses on workload balance. Our evaluations on different Transformer models demonstrate the capabilities of Galvatron-BMW in automating distributed training under varying GPU memory constraints. Across all tested scenarios, Galvatron-BMW consistently achieves superior system throughput, surpassing previous approaches that rely on limited parallelism strategies.
arXiv admin note: substantial text overlap with arXiv:2211.13878
References in corpus (10)
- Scaling Laws for Neural Language Models
- Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
- Supporting Very Large Models using Automatic Dataflow Graph Partitioning
- Do Transformers Really Perform Bad for Graph Representation?
- Galvatron: Efficient Transformer Training over Multiple GPUs Using Automatic Parallelism
- TensorOpt: Exploring the Tradeoffs in Distributed DNN Training with Auto-Parallelism
- Enabling Compute-Communication Overlap in Distributed Deep Learning Training Platforms
- GSPMD: General and Scalable Parallelization for ML Computation Graphs
- Automatic Cross-Replica Sharding of Weight Update in Data-Parallel Training
- HetuMoE: An Efficient Trillion-scale Mixture-of-Expert Distributed Training System
Cited by in corpus (4)
- Apt-Serve: Adaptive Request Scheduling on Hybrid Cache for Scalable LLM Inference Serving
- LobRA: Multi-tenant Fine-tuning over Heterogeneous Data
- Hydraulis: Balancing Large Transformer Model Training via Co-designing Parallel Strategies and Data Assignment
- MEMO: Fine-grained Tensor Management For Ultra-long Context LLM Training