3 papers
cs.DC2026
ScaleAcross Explorer: Exploring Communication Optimization for Scale-Across AI Model Training
Minghao Li, Alicia Golden, Samuel Hsia +14
The rapid scaling of large language model training requires distributing GPU resources across multiple data center buildings and regions. We refer to such paradigm as "scale-across…
cs.DC2026
Exploiting Multicast for Accelerating Collective Communication
Chao Xu, Xu Zhang, Zihang Luo +5
Reducing collective communication latency is a critical goal for large model training and inference in both academia and industry. Many-to-many communications, such as AllGather an…
cs.DC2024
Accelerating Heterogeneous Tensor Parallelism via Flexible Workload Control
Zhigang Wang, Xu Zhang, Ning Wang +5
Transformer-based models are becoming deeper and larger recently. For better scalability, an underlying training solution in industry is to split billions of parameters (tensors) i…