1 paper · 1 filter
Bowen Peng, Lizhang Chen, Baiyu Su +3
Scaling neural network training increasingly depends on synchronous data-parallelism, yet full-precision gradient all-reduce imposes a severe communication bottleneck. We propose D…