19 citations · 59 across the 9 of their papers we have counts for
1 paper · 1 filter
Wei Wen, Cong Xu, Feng Yan +4
High network communication cost for synchronizing gradients and parameters is the well-known bottleneck of distributed training. In this work, we propose TernGrad that uses ternary…