708 citations · 1.5k across the 7 of their papers we have counts for
1 paper · 2 filters
Yuanzhong Xu, HyoukJoong Lee, Dehao Chen +3
In data-parallel synchronous training of deep neural networks, different devices (replicas) run the same program with different partitions of the training batch, but weight update…