1 paper
Qinyi Luo, Jiaao He, Youwei Zhuo +1
Distributed deep learning training usually adopts All-Reduce as the synchronization mechanism for data parallel algorithms due to its high performance in homogeneous environment. H…