2 citations · 2 across the 1 of their papers we have counts for
1 paper · 1 filter
Xiaowu Dai, Yuhua Zhu
Stochastic gradient descent (SGD) is almost ubiquitously used for training non-convex optimization tasks. Recently, a hypothesis proposed by Keskar et al. [2017] that large batch m…