296 citations · 948 across the 18 of their papers we have counts for
22 papers
Data Scaling Laws in NMT: The Effect of Noise and Architecture
Yamini Bansal, Behrooz Ghorbani, Ankush Garg +5
In this work, we study the effect of varying the architecture and training data quality on the data scaling properties of Neural Machine Translation (NMT). First, we establish that…
A Loss Curvature Perspective on Training Instability in Deep Learning
Justin Gilmer, Behrooz Ghorbani, Ankush Garg +6
In this work, we study the evolution of the loss Hessian across many classification tasks in order to understand the effect the curvature of the loss has on the training dynamics.…
Exploring the Limits of Large Scale Pre-training
Samira Abnar, Mostafa Dehghani, Behnam Neyshabur +1
Recent developments in large-scale machine learning suggest that by scaling up data, model size and training time properly, one might observe that improvements in pre-training woul…
The Evolution of Out-of-Distribution Robustness Throughout Fine-Tuning
Anders Andreassen, Yasaman Bahri, Behnam Neyshabur +1
Although machine learning models typically experience a drop in performance on out-of-distribution data, accuracies on in- versus out-of-distribution data are widely observed to fo…
Deep Learning Through the Lens of Example Difficulty
Robert J. N. Baldock, Hartmut Maennel, Behnam Neyshabur
Existing work on understanding deep learning often employs measures that compress all data-dependent information into a few numbers. In this work, we adopt a perspective based on t…
When Do Curricula Work?
Xiaoxia Wu, Ethan Dyer, Behnam Neyshabur
Inspired by human learning, researchers have proposed ordering examples during training based on their difficulty. Both curriculum learning, exposing a network to easier examples e…