activity
20152022
most citedExploring Generalization in Deep Learning

296 citations · 948 across the 18 of their papers we have counts for

collaborators

22 papers

cs.LG20228 cited

Data Scaling Laws in NMT: The Effect of Noise and Architecture

Yamini Bansal, Behrooz Ghorbani, Ankush Garg +5

In this work, we study the effect of varying the architecture and training data quality on the data scaling properties of Neural Machine Translation (NMT). First, we establish that…

cs.LG20216 cited

A Loss Curvature Perspective on Training Instability in Deep Learning

Justin Gilmer, Behrooz Ghorbani, Ankush Garg +6

In this work, we study the evolution of the loss Hessian across many classification tasks in order to understand the effect the curvature of the loss has on the training dynamics.…

cs.LG202135 cited

Exploring the Limits of Large Scale Pre-training

Samira Abnar, Mostafa Dehghani, Behnam Neyshabur +1

Recent developments in large-scale machine learning suggest that by scaling up data, model size and training time properly, one might observe that improvements in pre-training woul…

cs.LG202120 cited

The Evolution of Out-of-Distribution Robustness Throughout Fine-Tuning

Anders Andreassen, Yasaman Bahri, Behnam Neyshabur +1

Although machine learning models typically experience a drop in performance on out-of-distribution data, accuracies on in- versus out-of-distribution data are widely observed to fo…

cs.LG202120 cited

Deep Learning Through the Lens of Example Difficulty

Robert J. N. Baldock, Hartmut Maennel, Behnam Neyshabur

Existing work on understanding deep learning often employs measures that compress all data-dependent information into a few numbers. In this work, we adopt a perspective based on t…

cs.LG20207 cited

When Do Curricula Work?

Xiaoxia Wu, Ethan Dyer, Behnam Neyshabur

Inspired by human learning, researchers have proposed ordering examples during training based on their difficulty. Both curriculum learning, exposing a network to easier examples e…