1 paper · 1 filter
Wei Wang, Nengneng Yu, Sixian Xiong +1
Modern ML training and inference now span tens to tens of thousands of GPUs, where network faults can waste 10--15\% of GPU hours due to slow recovery. Common network errors and li…