2 papers
cs.NI2026
On Topology's Role in ML Training Performance
Sarah McClure, Tegan Wilson, Brad Karp +4
Modern machine learning training workloads run on large-scale networks of compute accelerators. The networks commonly deployed in these systems are typically variations of two basi…
cs.LG2024
Beyond Throughput and Compression Ratios: Towards High End-to-end Utility of Gradient Compression
Wenchen Han, Shay Vargaftik, Michael Mitzenmacher +2
Gradient aggregation has long been identified as a major bottleneck in today's large-scale distributed machine learning training systems. One promising solution to mitigate such bo…