1 paper
Lichen Pan, Juncheng Liu, Yongquan Fu +4
Distributed deep neural network training necessitates efficient GPU collective communications, which are inherently susceptible to deadlocks. GPU collective deadlocks arise easily…