1 paper · 1 filter
Hailing Cheng, Tao Huang, Chen Zhu +1
Training large neural networks with data-parallel stochastic gradient descent allocates N GPU replicas to compute effectively identical updates -- a practice that leaves the rich s…