571 citations · 576 across the 2 of their papers we have counts for
2 papers
cs.LG2016★ 571 cited
On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal +2
The stochastic gradient descent (SGD) method and its variants are algorithms of choice for many Deep Learning tasks. These methods operate in a small-batch regime wherein a fractio…
math.NA2014★ 5 cited
Zolotarev Quadrature Rules and Load Balancing for the FEAST Eigensolver
Stefan Guettel, Eric Polizzi, Ping Tak Peter Tang +1
The FEAST method for solving large sparse eigenproblems is equivalent to subspace iteration with an approximate spectral projector and implicit orthogonalization. This relation all…