571 citations · 571 across the 2 of their papers we have counts for
2 papers
cs.LG2016★ 571 cited
On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal +2
The stochastic gradient descent (SGD) method and its variants are algorithms of choice for many Deep Learning tasks. These methods operate in a small-batch regime wherein a fractio…
math.OC2014
An Algorithm for Quadratic -Regularized Optimization with a Flexible Active-Set Strategy
Stefan Solntsev, Jorge Nocedal, Richard Byrd
We present an active-set method for minimizing an objective that is the sum of a convex quadratic and regularization term. Unlike two-phase methods that combine a first-or…