1 paper · 1 filter
Zhuang Wang, Xinyu Wu, T. S. Eugene Ng
Large-scale distributed training is increasingly becoming communication bound. Many gradient compression algorithms have been proposed to reduce the communication overhead and impr…