97 citations · 115 across the 3 of their papers we have counts for
11 papers
Proving Test Set Contamination in Black Box Language Models
Yonatan Oren, Nicole Meister, Niladri Chatterji +2
Large language models are trained on vast amounts of internet data, prompting concerns and speculation that they have memorized public benchmarks. Going from speculation to proof o…
When does gradient descent with logistic loss interpolate using deep networks with smoothed ReLU activations?
Niladri S. Chatterji, Philip M. Long, Peter L. Bartlett
We establish conditions under which gradient descent applied to fixed-width deep networks drives the logistic loss to zero, and prove bounds on the rate of convergence. Our analysi…
When does gradient descent with logistic loss find interpolating two-layer networks?
Niladri S. Chatterji, Philip M. Long, Peter L. Bartlett
We study the training of finite-width two-layer smoothed ReLU networks for binary classification using the logistic loss. We show that gradient descent drives the training loss to…
The intriguing role of module criticality in the generalization of deep networks
Niladri S. Chatterji, Behnam Neyshabur, Hanie Sedghi
We study the phenomenon that some modules of deep neural networks (DNNs) are more critical than others. Meaning that rewinding their parameter values back to initialization, while…
Langevin Monte Carlo without smoothness
Niladri S. Chatterji, Jelena Diakonikolas, Michael I. Jordan +1
Langevin Monte Carlo (LMC) is an iterative algorithm used to generate samples from a distribution that is known only up to a normalizing constant. The nonasymptotic dependence of i…
OSOM: A simultaneously optimal algorithm for multi-armed and linear contextual bandits
Niladri S. Chatterji, Vidya Muthukumar, Peter L. Bartlett
We consider the stochastic linear (multi-armed) contextual bandit problem with the possibility of hidden simple multi-armed bandit structure in which the rewards are independent of…