3 papers
cs.DC2026
ScaleAcross Explorer: Exploring Communication Optimization for Scale-Across AI Model Training
Minghao Li, Alicia Golden, Samuel Hsia +14
The rapid scaling of large language model training requires distributing GPU resources across multiple data center buildings and regions. We refer to such paradigm as "scale-across…
cs.DC2026
PRISM: Probabilistic Runtime Insights and Scalable Performance Modeling for Large-Scale Distributed Training
Alicia Golden, Michael Kuchnik, Samuel Hsia +4
Large model training beyond tens of thousands of GPUs is an uncharted territory. At such scales, disruptions to the training process are not a matter of if, but a matter of when --…
cs.DC2025
Revisiting Reliability in Large-Scale Machine Learning Research Clusters
Apostolos Kokolis, Michael Kuchnik, John Hoffman +7
Reliability is a fundamental challenge in operating large-scale machine learning (ML) infrastructures, particularly as the scale of ML models and training clusters continues to gro…