2 papers
cs.DC2025
Oases: Efficient Large-Scale Model Training on Commodity Servers via Overlapped and Automated Tensor Model Parallelism
Shengwei Li, Zhiquan Lai, Dongsheng Li +5
Deep learning is experiencing a rise in large-scale models. Training large-scale models is costly, prompting researchers to train large-scale models on commodity servers that more…
cs.DC2024
Pro-Prophet: A Systematic Load Balancing Method for Efficient Parallel Training of Large-scale MoE Models
Wei Wang, Zhiquan Lai, Shengwei Li +5
The size of deep learning models has been increasing to enhance model quality. The linear increase in training computation budget with model size means that training an extremely l…