5 papers
PCCL: Photonic circuit-switched collective communication for distributed ML
Abhishek Vijaya Kumar, Arjun Devraj, Rachee Singh
Modern distributed ML suffers from a fundamental gap between the theoretical and realized performance of collective communication algorithms due to congestion and hop-count induced…
Morphlux: Transforming Torus Fabrics for Efficient Multi-tenant ML
Abhishek Vijaya Kumar, Eric Ding, Arjun Devraj +2
We develop Morphlux, a server-scale programmable photonic fabric to interconnect accelerators within servers. We show that augmenting state-of-the-art torus-based ML data-centers w…
Efficient AllReduce with Stragglers
Arjun Devraj, Eric Ding, Abhishek Vijaya Kumar +2
Distributed machine learning workloads use data and tensor parallelism for training and inference, both of which rely on the AllReduce collective to synchronize gradients or activa…
LUMION: Fast Fault Recovery for ML Jobs Using Programmable Optical Fabrics
Abhishek Vijaya Kumar, Eric Ding, Arjun Devraj +2
When accelerators fail in modern ML datacenters, operators migrate the affected ML training or inference jobs to entirely new racks. This approach, while preserving network perform…
Chip-to-chip photonic connectivity in multi-accelerator servers for ML
Abhishek Vijaya Kumar, Arjun Devraj, Darius Bunandar +1
We present a rack-scale compute architecture for ML using multi-accelerator servers connected via chip-to-chip silicon photonic components. Our architecture achieves (1) multi-tena…