papers

Publications (5)

cs.NI2026

Resilient AI Supercomputer Networking using MRC and SRv6

Joao Araujo, Alex Chow, Mark Handley +47

Tail latency dominates the performance of synchronous pretraining jobs when running at very large scales. We describe a three-pronged approach: (1) a new RDMA-based transport proto…

cs.NI2026

In-Network Collective Operations: Game Changer or Challenge for AI Workloads?

Torsten Hoefler, Mikhail Khalilov, Josiah Clark +9

This paper summarizes the opportunities of in-network collective operations (INC) for accelerated collective operations in AI workloads. We provide sufficient detail to make this i…

cs.DC2026

Demystifying NCCL: An In-depth Analysis of GPU Communication Protocols and Algorithms

Zhiyi Hu, Siyuan Shen, Tommaso Bonato +6

The NVIDIA Collective Communication Library (NCCL) is a critical software layer enabling high-performance collectives on large-scale GPU clusters. Despite being open source with a…

cs.NI2026

The Multipath Reliable Connection (MRC) Transport

Rip Sohan, Eric Spada, Eric Davis +36

MRC is an open, production-grade transport designed for large-scale AI/ML training over best-effort Ethernet. It extends RoCEv2 with explicit, composable primitives for per-packet…

cs.NI2025

Ultra Ethernet's Design Principles and Architectural Innovations

Torsten Hoefler, Karen Schramm, Eric Spada +13

The recently released Ultra Ethernet (UE) 1.0 specification defines a transformative High-Performance Ethernet standard for future Artificial Intelligence (AI) and High-Performance…