7 papers
CCL-Bench 1.0: A Trace-Based Benchmark for LLM Infrastructure
Eric Ding, Byungsoo Oh, Bhaskar Kataria +8
Evaluative claims about LLM infrastructure -- ``workload X is fastest on hardware Y with software Z'' -- depend on a complex configuration space spanning hardware accelerators, int…
Stable and Fault-Tolerant Decentralized Traffic Engineering
Arjun Devraj, Umesh Krishnaswamy, Ying Zhang +4
Cloud providers have recently decentralized their wide-area network traffic engineering (TE) systems to contain the impact of TE controller failures. In the decentralized design, a…
LUMION: Fast Fault Recovery for ML Jobs Using Programmable Optical Fabrics
Abhishek Vijaya Kumar, Eric Ding, Arjun Devraj +2
When accelerators fail in modern ML datacenters, operators migrate the affected ML training or inference jobs to entirely new racks. This approach, while preserving network perform…
Morphlux: Transforming Torus Fabrics for Efficient Multi-tenant ML
Abhishek Vijaya Kumar, Eric Ding, Arjun Devraj +2
We develop Morphlux, a server-scale programmable photonic fabric to interconnect accelerators within servers. We show that augmenting state-of-the-art torus-based ML data-centers w…
Efficient AllReduce with Stragglers
Arjun Devraj, Eric Ding, Abhishek Vijaya Kumar +2
Distributed machine learning workloads use data and tensor parallelism for training and inference, both of which rely on the AllReduce collective to synchronize gradients or activa…
PCCL: Photonic circuit-switched collective communication for distributed ML
Abhishek Vijaya Kumar, Arjun Devraj, Rachee Singh
Modern distributed ML suffers from a fundamental gap between the theoretical and realized performance of collective communication algorithms due to congestion and hop-count induced…