7 citations · 9 across the 4 of their papers we have counts for
4 papers
In-Network Collective Operations: Game Changer or Challenge for AI Workloads?
Torsten Hoefler, Mikhail Khalilov, Josiah Clark +9
This paper summarizes the opportunities of in-network collective operations (INC) for accelerated collective operations in AI workloads. We provide sufficient detail to make this i…
Uno: A One-Stop Solution for Inter- and Intra-Datacenter Congestion Control and Reliable Connectivity
Tommaso Bonato, Sepehr Abdous, Abdul Kabbani +11
Cloud computing and AI workloads are driving unprecedented demand for efficient communication within and across datacenters. However, the coexistence of intra- and inter-datacenter…
Ultra Ethernet's Design Principles and Architectural Innovations
Torsten Hoefler, Karen Schramm, Eric Spada +13
The recently released Ultra Ethernet (UE) 1.0 specification defines a transformative High-Performance Ethernet standard for future Artificial Intelligence (AI) and High-Performance…
SDR-RDMA: Software-Defined Reliability Architecture for Planetary Scale RDMA Communication
Mikhail Khalilov, Siyuan Shen, Marcin Chrapek +16
RDMA is vital for efficient distributed training across datacenters, but millisecond-scale latencies complicate the design of its reliability layer. We show that depending on long-…