2 citations · 2 across the 1 of their papers we have counts for
2 papers
cs.LG2020★ 2 cited
Scalable and Practical Natural Gradient for Large-Scale Deep Learning
Kazuki Osawa, Yohei Tsuji, Yuichiro Ueno +3
Large-scale distributed training of deep neural networks results in models with worse generalization performance as a result of the increase in the effective mini-batch size. Previ…
cs.LG2018
Large-Scale Distributed Second-Order Optimization Using Kronecker-Factored Approximate Curvature for Deep Convolutional Neural Networks
Kazuki Osawa, Yohei Tsuji, Yuichiro Ueno +3
Large-scale distributed training of deep neural networks suffer from the generalization gap caused by the increase in the effective mini-batch size. Previous approaches try to solv…