9 papers
Opus: Photonic Rail-Optimized Fabric in ML Datacenters
Eric Ding, Barry Lyu, Bhaskar Kataria +1
Rail-optimized network fabrics have become the de facto datacenter scale-out fabric for large-scale ML training. However, the use of high-radix electrical switches to provide all-t…
CCL-Bench 1.0: A Trace-Based Benchmark for LLM Infrastructure
Eric Ding, Byungsoo Oh, Bhaskar Kataria +8
Evaluative claims about LLM infrastructure -- ``workload X is fastest on hardware Y with software Z'' -- depend on a complex configuration space spanning hardware accelerators, int…
Propius: A Platform for Collaborative Machine Learning across the Edge and the Cloud
Eric Ding
Collaborative Machine Learning is a paradigm in the field of distributed machine learning, designed to address the challenges of data privacy, communication overhead, and model het…
LUMION: Fast Fault Recovery for ML Jobs Using Programmable Optical Fabrics
Abhishek Vijaya Kumar, Eric Ding, Arjun Devraj +2
When accelerators fail in modern ML datacenters, operators migrate the affected ML training or inference jobs to entirely new racks. This approach, while preserving network perform…
Morphlux: Transforming Torus Fabrics for Efficient Multi-tenant ML
Abhishek Vijaya Kumar, Eric Ding, Arjun Devraj +2
We develop Morphlux, a server-scale programmable photonic fabric to interconnect accelerators within servers. We show that augmenting state-of-the-art torus-based ML data-centers w…
Efficient AllReduce with Stragglers
Arjun Devraj, Eric Ding, Abhishek Vijaya Kumar +2
Distributed machine learning workloads use data and tensor parallelism for training and inference, both of which rely on the AllReduce collective to synchronize gradients or activa…