2 papers
cs.NI2025
SHIFT: Exploring the Boundary of RDMA Network Fault Tolerance
Shengkai Lin, Kairui Zhou, Hongtao Zhang +7
Under gang scheduling for large-scale distributed large language model (LLM) training, a single network anomaly can stall or abort an entire job. Current network fault tolerance me…
cs.DC2023
Efficient All-to-All Collective Communication Schedules for Direct-Connect Topologies
Prithwish Basu, Liangyu Zhao, Jason Fantl +3
The all-to-all collective communications primitive is widely used in machine learning (ML) and high performance computing (HPC) workloads, and optimizing its performance is of inte…