1 citations · 1 across the 3 of their papers we have counts for
4 papers
SDR-RDMA: Software-Defined Reliability Architecture for Planetary Scale RDMA Communication
Mikhail Khalilov, Siyuan Shen, Marcin Chrapek +16
RDMA is vital for efficient distributed training across datacenters, but millisecond-scale latencies complicate the design of its reliability layer. We show that depending on long-…
Noise in the Clouds: Influence of Network Performance Variability on Application Scalability
Daniele De Sensi, Tiziano De Matteis, Konstantin Taranov +3
Cloud computing represents an appealing opportunity for cost-effective deployment of HPC workloads on the best-fitting hardware. However, although cloud and on-premise HPC systems…
HammingMesh: A Network Topology for Large-Scale Deep Learning
Torsten Hoefler, Tommaso Bonato, Daniele De Sensi +7
Numerous microarchitectural optimizations unlocked tremendous processing power for deep neural networks that in turn fueled the AI revolution. With the exhaustion of such optimizat…
High-Performance Routing with Multipathing and Path Diversity in Ethernet and HPC Networks
Maciej Besta, Jens Domke, Marcel Schneider +5
The recent line of research into topology design focuses on lowering network diameter. Many low-diameter topologies such as Slim Fly or Jellyfish that substantially reduce cost, po…