26 citations · 37 across the 6 of their papers we have counts for
1 paper · 1 filter
Kazuki Osawa, Yohei Tsuji, Yuichiro Ueno +3
Large-scale distributed training of deep neural networks suffer from the generalization gap caused by the increase in the effective mini-batch size. Previous approaches try to solv…