collaborators

7 papers

cs.DC2026

CCL-Bench 1.0: A Trace-Based Benchmark for LLM Infrastructure

Eric Ding, Byungsoo Oh, Bhaskar Kataria +8

Evaluative claims about LLM infrastructure -- ``workload X is fastest on hardware Y with software Z'' -- depend on a complex configuration space spanning hardware accelerators, int…

cs.NI2025

Stable and Fault-Tolerant Decentralized Traffic Engineering

Arjun Devraj, Umesh Krishnaswamy, Ying Zhang +4

Cloud providers have recently decentralized their wide-area network traffic engineering (TE) systems to contain the impact of TE controller failures. In the decentralized design, a…

cs.LG2025

LUMION: Fast Fault Recovery for ML Jobs Using Programmable Optical Fabrics

Abhishek Vijaya Kumar, Eric Ding, Arjun Devraj +2

When accelerators fail in modern ML datacenters, operators migrate the affected ML training or inference jobs to entirely new racks. This approach, while preserving network perform…

cs.NI2025

Morphlux: Transforming Torus Fabrics for Efficient Multi-tenant ML

Abhishek Vijaya Kumar, Eric Ding, Arjun Devraj +2

We develop Morphlux, a server-scale programmable photonic fabric to interconnect accelerators within servers. We show that augmenting state-of-the-art torus-based ML data-centers w…

cs.LG2025

Efficient AllReduce with Stragglers

Arjun Devraj, Eric Ding, Abhishek Vijaya Kumar +2

Distributed machine learning workloads use data and tensor parallelism for training and inference, both of which rely on the AllReduce collective to synchronize gradients or activa…

cs.DC2025

PCCL: Photonic circuit-switched collective communication for distributed ML

Abhishek Vijaya Kumar, Arjun Devraj, Rachee Singh

Modern distributed ML suffers from a fundamental gap between the theoretical and realized performance of collective communication algorithms due to congestion and hop-count induced…