Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Efficient AllReduce with Stragglers
Arjun Devraj, Eric Ding, Abhishek Vijaya Kumar +2
Distributed machine learning workloads use data and tensor parallelism for training and inference, both of which rely on the AllReduce collective to synchronize gradients or activa…
cs.LG2025
LUMION: Fast Fault Recovery for ML Jobs Using Programmable Optical Fabrics
Abhishek Vijaya Kumar, Eric Ding, Arjun Devraj +2
When accelerators fail in modern ML datacenters, operators migrate the affected ML training or inference jobs to entirely new racks. This approach, while preserving network perform…