4 papers
MoX: Efficient MoE Routing on Direct-Connect Topologies
Ori Cohen, Jakob Krebs, Daniel Amir +1
Optically switched networks suit the regular communication of dense ML models, but MoE introduces sparse, runtime-dependent traffic. We show that efficient offline-optimized routin…
High-speed Networking for Giga-Scale AI Factories
Sajy Khashab, Albert Gran Alcoz, Alon Gal +11
As distributed model training scales to span hundreds of thousands of GPUs, scale-out networks face unprecedented performance and efficiency demands. NVIDIA Spectrum-X Ethernet has…
SprayCheck: Finding Gray Failures in Adaptive Routing Networks
Jakob Krebs, Daniel Amir, Shir Landau Feibish +1
Distributed machine learning (ML) training has become a dominant workload in modern data center networks, operating at massive scale with clusters comprising tens to hundreds of th…
ACOS: Arrays of Cheap Optical Switches
Daniel Amir, Ori Cohen, Jakob Krebs +1
Machine learning training places immense demands on cluster networks, motivating specialized architectures and co-design with parallelization strategies. Recent designs incorporati…