8 citations · 8 across the 1 of their papers we have counts for
4 papers
EROICA: Online Performance Troubleshooting for Large-scale Model Training
Yu Guan, Zhiyu Yin, Haoyu Chen +11
Troubleshooting performance problems of large model training (LMT) is immensely challenging, due to unprecedented scales of modern GPU clusters, the complexity of software-hardware…
TrainMover: An Interruption-Resilient Runtime for ML Training
ChonLam Lao, Jiaqi Gao, Jiamin Cao +13
Large-scale ML training jobs are frequently interrupted by hardware and software anomalies, failures, and management events. Existing solutions like checkpoint-restart or runtime r…
Enhancing Large-Scale AI Training Efficiency: The C4 Solution for Real-Time Anomaly Detection and Communication Optimization
Jianbo Dong, Bin Luo, Jun Zhang +22
The emergence of Large Language Models (LLMs) has necessitated the adoption of distributed training techniques, involving the deployment of thousands of GPUs to train a single mode…
DxPU: Large Scale Disaggregated GPU Pools in the Datacenter
Bowen He, Xiao Zheng, Yuan Chen +9
The rapid adoption of AI and convenience offered by cloud services have resulted in the growing demands for GPUs in the cloud. Generally, GPUs are physically attached to host serve…