collaborators

7 papers

cs.NI2026

The Multipath Reliable Connection (MRC) Transport

Rip Sohan, Eric Spada, Eric Davis +36

MRC is an open, production-grade transport designed for large-scale AI/ML training over best-effort Ethernet. It extends RoCEv2 with explicit, composable primitives for per-packet…

cs.NI2026

Resilient AI Supercomputer Networking using MRC and SRv6

Joao Araujo, Alex Chow, Mark Handley +47

Tail latency dominates the performance of synchronous pretraining jobs when running at very large scales. We describe a three-pronged approach: (1) a new RDMA-based transport proto…

cs.NI2026

SMaRTT: Sender-based Marked Rapidly-adapting Trimmed & Timed Transport

Tommaso Bonato, Abdul Kabbani, Ahmad Ghalayini +10

With the rapid growth of artificial intelligence (AI) workloads in datacenters, the Ultra Ethernet Consortium (UEC) has defined a new high-performance transport layer to deliver th…

cs.DC2026

UCCL-EP: Portable Expert-Parallel Communication

Ziming Mao, Yihan Zhang, Chihan Cui +9

Mixture-of-Experts (MoE) workloads rely on expert parallelism (EP) to achieve high GPU efficiency. State-of-the-art EP communication systems such as DeepEP demonstrate strong perfo…

cs.NI2025

Ultra Ethernet's Design Principles and Architectural Innovations

Torsten Hoefler, Karen Schramm, Eric Spada +13

The recently released Ultra Ethernet (UE) 1.0 specification defines a transformative High-Performance Ethernet standard for future Artificial Intelligence (AI) and High-Performance…

cs.NI2025

An Extensible Software Transport Layer for GPU Networking

Yang Zhou, Zhongjie Chen, Ziming Mao +11

Fast-evolving machine learning (ML) workloads have increasing requirements for networking. However, host network transport on RDMA NICs is hard to evolve, causing problems for ML w…