1 paper
Hiroaki Mikami, Hisahiro Suganuma, Pongsakorn U-chupala +2
Scaling the distributed deep learning to a massive GPU cluster level is challenging due to the instability of the large mini-batch training and the overhead of the gradient synchro…