collaborators

9 papers

cs.NI2026

Opus: Photonic Rail-Optimized Fabric in ML Datacenters

Eric Ding, Barry Lyu, Bhaskar Kataria +1

Rail-optimized network fabrics have become the de facto datacenter scale-out fabric for large-scale ML training. However, the use of high-radix electrical switches to provide all-t…

cs.DC2026

CCL-Bench 1.0: A Trace-Based Benchmark for LLM Infrastructure

Eric Ding, Byungsoo Oh, Bhaskar Kataria +8

Evaluative claims about LLM infrastructure -- ``workload X is fastest on hardware Y with software Z'' -- depend on a complex configuration space spanning hardware accelerators, int…

cs.DC2025

Propius: A Platform for Collaborative Machine Learning across the Edge and the Cloud

Eric Ding

Collaborative Machine Learning is a paradigm in the field of distributed machine learning, designed to address the challenges of data privacy, communication overhead, and model het…

cs.LG2025

LUMION: Fast Fault Recovery for ML Jobs Using Programmable Optical Fabrics

Abhishek Vijaya Kumar, Eric Ding, Arjun Devraj +2

When accelerators fail in modern ML datacenters, operators migrate the affected ML training or inference jobs to entirely new racks. This approach, while preserving network perform…

cs.NI2025

Morphlux: Transforming Torus Fabrics for Efficient Multi-tenant ML

Abhishek Vijaya Kumar, Eric Ding, Arjun Devraj +2

We develop Morphlux, a server-scale programmable photonic fabric to interconnect accelerators within servers. We show that augmenting state-of-the-art torus-based ML data-centers w…

cs.LG2025

Efficient AllReduce with Stragglers

Arjun Devraj, Eric Ding, Abhishek Vijaya Kumar +2

Distributed machine learning workloads use data and tensor parallelism for training and inference, both of which rely on the AllReduce collective to synchronize gradients or activa…