353 citations · 460 across the 9 of their papers we have counts for
Showing stat.MLShow all
2 papers · 1 filter
stat.ML2018
On the Relation Between the Sharpest Directions of DNN Loss and the SGD Step Length
Stanisław Jastrzębski, Zachary Kenton, Nicolas Ballas +3
Stochastic Gradient Descent (SGD) based training of neural networks with a large learning rate or a small batch-size typically ends in well-generalizing, flat regions of the weight…
stat.ML2017★ 353 cited
A Closer Look at Memorization in Deep Networks
Devansh Arpit, Stanisław Jastrzębski, Nicolas Ballas +8
We examine the role of memorization in deep learning, drawing connections to capacity, generalization, and adversarial robustness. While deep networks are capable of memorizing noi…