35 citations · 83 across the 8 of their papers we have counts for
13 papers
Leveraging Unlabeled Data to Track Memorization
Mahsa Forouzesh, Hanie Sedghi, Patrick Thiran
Deep neural networks may easily memorize noisy labels present in real-world data, which degrades their ability to generalize. It is therefore important to track and evaluate the ro…
Layer-Stack Temperature Scaling
Amr Khalifa, Michael C. Mozer, Hanie Sedghi +2
Recent works demonstrate that early layers in a neural network contain useful information for prediction. Inspired by this, we show that extending temperature scaling across all la…
Teaching Algorithmic Reasoning via In-context Learning
Hattie Zhou, Azade Nova, Hugo Larochelle +3
Large language models (LLMs) have shown increasing in-context learning capabilities through scaling up model and data size. Despite this progress, LLMs are still unable to solve al…
Exploring the Limits of Large Scale Pre-training
Samira Abnar, Mostafa Dehghani, Behnam Neyshabur +1
Recent developments in large-scale machine learning suggest that by scaling up data, model size and training time properly, one might observe that improvements in pre-training woul…
Gradual Domain Adaptation in the Wild:When Intermediate Distributions are Absent
Samira Abnar, Rianne van den Berg, Golnaz Ghiasi +3
We focus on the problem of domain adaptation when the goal is shifting the model towards the target distribution, rather than learning domain invariant representations. It has been…
The Deep Bootstrap Framework: Good Online Learners are Good Offline Generalizers
Preetum Nakkiran, Behnam Neyshabur, Hanie Sedghi
We propose a new framework for reasoning about generalization in deep learning. The core idea is to couple the Real World, where optimizers take stochastic gradient steps on the em…