activity
20152022
most citedExploring the Limits of Large Scale Pre-training

35 citations · 83 across the 8 of their papers we have counts for

collaborators

13 papers

cs.LG2022

Leveraging Unlabeled Data to Track Memorization

Mahsa Forouzesh, Hanie Sedghi, Patrick Thiran

Deep neural networks may easily memorize noisy labels present in real-world data, which degrades their ability to generalize. It is therefore important to track and evaluate the ro…

cs.LG2022

Layer-Stack Temperature Scaling

Amr Khalifa, Michael C. Mozer, Hanie Sedghi +2

Recent works demonstrate that early layers in a neural network contain useful information for prediction. Inspired by this, we show that extending temperature scaling across all la…

cs.LG202224 cited

Teaching Algorithmic Reasoning via In-context Learning

Hattie Zhou, Azade Nova, Hugo Larochelle +3

Large language models (LLMs) have shown increasing in-context learning capabilities through scaling up model and data size. Despite this progress, LLMs are still unable to solve al…

cs.LG202135 cited

Exploring the Limits of Large Scale Pre-training

Samira Abnar, Mostafa Dehghani, Behnam Neyshabur +1

Recent developments in large-scale machine learning suggest that by scaling up data, model size and training time properly, one might observe that improvements in pre-training woul…

cs.LG20214 cited

Gradual Domain Adaptation in the Wild:When Intermediate Distributions are Absent

Samira Abnar, Rianne van den Berg, Golnaz Ghiasi +3

We focus on the problem of domain adaptation when the goal is shifting the model towards the target distribution, rather than learning domain invariant representations. It has been…

cs.LG2020

The Deep Bootstrap Framework: Good Online Learners are Good Offline Generalizers

Preetum Nakkiran, Behnam Neyshabur, Hanie Sedghi

We propose a new framework for reasoning about generalization in deep learning. The core idea is to couple the Real World, where optimizers take stochastic gradient steps on the em…