3 citations · 3 across the 1 of their papers we have counts for
1 paper
Kosuke Haruki, Taiji Suzuki, Yohei Hamakawa +4
Large-batch stochastic gradient descent (SGD) is widely used for training in distributed deep learning because of its training-time efficiency, however, extremely large-batch SGD l…