97 citations · 116 across the 10 of their papers we have counts for
9 papers · 1 filter
When does gradient descent with logistic loss interpolate using deep networks with smoothed ReLU activations?
Niladri S. Chatterji, Philip M. Long, Peter L. Bartlett
We establish conditions under which gradient descent applied to fixed-width deep networks drives the logistic loss to zero, and prove bounds on the rate of convergence. Our analysi…
When does gradient descent with logistic loss find interpolating two-layer networks?
Niladri S. Chatterji, Philip M. Long, Peter L. Bartlett
We study the training of finite-width two-layer smoothed ReLU networks for binary classification using the logistic loss. We show that gradient descent drives the training loss to…
Langevin Monte Carlo without smoothness
Niladri S. Chatterji, Jelena Diakonikolas, Michael I. Jordan +1
Langevin Monte Carlo (LMC) is an iterative algorithm used to generate samples from a distribution that is known only up to a normalizing constant. The nonasymptotic dependence of i…
OSOM: A simultaneously optimal algorithm for multi-armed and linear contextual bandits
Niladri S. Chatterji, Vidya Muthukumar, Peter L. Bartlett
We consider the stochastic linear (multi-armed) contextual bandit problem with the possibility of hidden simple multi-armed bandit structure in which the rewards are independent of…
Is There an Analog of Nesterov Acceleration for MCMC?
Yi-An Ma, Niladri Chatterji, Xiang Cheng +3
We formulate gradient-based Markov chain Monte Carlo (MCMC) sampling as optimization on the space of probability measures, with Kullback-Leibler (KL) divergence as the objective fu…
Sharp convergence rates for Langevin dynamics in the nonconvex setting
Xiang Cheng, Niladri S. Chatterji, Yasin Abbasi-Yadkori +2
We study the problem of sampling from a distribution , where the function is -smooth everywhere and -strongly convex outside a ball…