128 citations · 273 across the 23 of their papers we have counts for
3 papers · 1 filter
When does gradient descent with logistic loss find interpolating two-layer networks?
Niladri S. Chatterji, Philip M. Long, Peter L. Bartlett
We study the training of finite-width two-layer smoothed ReLU networks for binary classification using the logistic loss. We show that gradient descent drives the training loss to…
Finite-sample Analysis of Interpolating Linear Classifiers in the Overparameterized Regime
Niladri S. Chatterji, Philip M. Long
We prove bounds on the population risk of the maximum margin algorithm for two-class linear classification. For linearly separable training data, the maximum margin algorithm has b…
Oracle Lower Bounds for Stochastic Gradient Sampling Algorithms
Niladri S. Chatterji, Peter L. Bartlett, Philip M. Long
We consider the problem of sampling from a strongly log-concave density in , and prove an information theoretic lower bound on the number of stochastic gradient queri…