2 papers
cs.DC2026
Comprehensive Deadlock Prevention for GPU Collective Communication
Lichen Pan, Juncheng Liu, Yongquan Fu +4
Distributed deep neural network training necessitates efficient GPU collective communications, which are inherently susceptible to deadlocks. GPU collective deadlocks arise easily…
cs.DC2025
AutoDDL: Automatic Distributed Deep Learning with Near-Optimal Bandwidth Cost
Jinfan Chen, Shigang Li, Ran Gun +2
Recent advances in deep learning are driven by the growing scale of computation, data, and models. However, efficiently training large-scale models on distributed systems requires…