3 papers
cs.NI2026
High-speed Networking for Giga-Scale AI Factories
Sajy Khashab, Albert Gran Alcoz, Alon Gal +11
As distributed model training scales to span hundreds of thousands of GPUs, scale-out networks face unprecedented performance and efficiency demands. NVIDIA Spectrum-X Ethernet has…
cs.NI2026
Avoiding Cross-Datacenter Collective Congestion via Disaggregated Buffering
Mariano Scazzariello, Noga H. Rotman, Dima Gavrilenko +6
LLM training at the scale of tens of thousands of GPUs now spans multiple datacenters (DC), making cross-DC collectives over long-haul links unavoidable. A critical and overlooked…
cs.NI2020
LB Scalability: Achieving the Right Balance Between Being Stateful and Stateless
Reuven Cohen, Matty Kadosh, Alan Lo +1
A high performance Layer-4 load balancer (LB) is one of the most important components of a cloud service infrastructure. Such an LB uses network and transport layer information for…