3 papers
cs.AR2025
RailX: A Flexible, Scalable, and Low-Cost Network Architecture for Hyper-Scale LLM Training Systems
Yinxiao Feng, Tiancheng Chen, Yuchen Wei +5
Increasingly large AI workloads are calling for hyper-scale infrastructure; however, traditional interconnection network architecture is neither scalable nor cost-effective enough.…
cs.DC2025
ATLAHS: An Application-centric Network Simulator Toolchain for AI, HPC, and Distributed Storage
Siyuan Shen, Tommaso Bonato, Zhiyi Hu +3
Network simulators play a crucial role in evaluating the performance of large-scale systems. However, existing simulators rely heavily on synthetic microbenchmarks or narrowly focu…
cs.NI2025
SDR-RDMA: Software-Defined Reliability Architecture for Planetary Scale RDMA Communication
Mikhail Khalilov, Siyuan Shen, Marcin Chrapek +16
RDMA is vital for efficient distributed training across datacenters, but millisecond-scale latencies complicate the design of its reliability layer. We show that depending on long-…