7 papers
The Multipath Reliable Connection (MRC) Transport
Rip Sohan, Eric Spada, Eric Davis +36
MRC is an open, production-grade transport designed for large-scale AI/ML training over best-effort Ethernet. It extends RoCEv2 with explicit, composable primitives for per-packet…
Resilient AI Supercomputer Networking using MRC and SRv6
Joao Araujo, Alex Chow, Mark Handley +47
Tail latency dominates the performance of synchronous pretraining jobs when running at very large scales. We describe a three-pronged approach: (1) a new RDMA-based transport proto…
SMaRTT: Sender-based Marked Rapidly-adapting Trimmed & Timed Transport
Tommaso Bonato, Abdul Kabbani, Ahmad Ghalayini +10
With the rapid growth of artificial intelligence (AI) workloads in datacenters, the Ultra Ethernet Consortium (UEC) has defined a new high-performance transport layer to deliver th…
UCCL-EP: Portable Expert-Parallel Communication
Ziming Mao, Yihan Zhang, Chihan Cui +9
Mixture-of-Experts (MoE) workloads rely on expert parallelism (EP) to achieve high GPU efficiency. State-of-the-art EP communication systems such as DeepEP demonstrate strong perfo…
Ultra Ethernet's Design Principles and Architectural Innovations
Torsten Hoefler, Karen Schramm, Eric Spada +13
The recently released Ultra Ethernet (UE) 1.0 specification defines a transformative High-Performance Ethernet standard for future Artificial Intelligence (AI) and High-Performance…
An Extensible Software Transport Layer for GPU Networking
Yang Zhou, Zhongjie Chen, Ziming Mao +11
Fast-evolving machine learning (ML) workloads have increasing requirements for networking. However, host network transport on RDMA NICs is hard to evolve, causing problems for ML w…