collaborators

5 papers

cs.NI2026

The Multipath Reliable Connection (MRC) Transport

Rip Sohan, Eric Spada, Eric Davis +36

MRC is an open, production-grade transport designed for large-scale AI/ML training over best-effort Ethernet. It extends RoCEv2 with explicit, composable primitives for per-packet…

cs.NI2026

Resilient AI Supercomputer Networking using MRC and SRv6

Joao Araujo, Alex Chow, Mark Handley +47

Tail latency dominates the performance of synchronous pretraining jobs when running at very large scales. We describe a three-pronged approach: (1) a new RDMA-based transport proto…

cs.DC2026

Demystifying NCCL: An In-depth Analysis of GPU Communication Protocols and Algorithms

Zhiyi Hu, Siyuan Shen, Tommaso Bonato +6

The NVIDIA Collective Communication Library (NCCL) is a critical software layer enabling high-performance collectives on large-scale GPU clusters. Despite being open source with a…

cs.NI2026

In-Network Collective Operations: Game Changer or Challenge for AI Workloads?

Torsten Hoefler, Mikhail Khalilov, Josiah Clark +9

This paper summarizes the opportunities of in-network collective operations (INC) for accelerated collective operations in AI workloads. We provide sufficient detail to make this i…

cs.NI2025

Ultra Ethernet's Design Principles and Architectural Innovations

Torsten Hoefler, Karen Schramm, Eric Spada +13

The recently released Ultra Ethernet (UE) 1.0 specification defines a transformative High-Performance Ethernet standard for future Artificial Intelligence (AI) and High-Performance…