174 citations · 612 across the 29 of their papers we have counts for
18 papers · 1 filter
When does gradient descent with logistic loss interpolate using deep networks with smoothed ReLU activations?
Niladri S. Chatterji, Philip M. Long, Peter L. Bartlett
We establish conditions under which gradient descent applied to fixed-width deep networks drives the logistic loss to zero, and prove bounds on the rate of convergence. Our analysi…
When does gradient descent with logistic loss find interpolating two-layer networks?
Niladri S. Chatterji, Philip M. Long, Peter L. Bartlett
We study the training of finite-width two-layer smoothed ReLU networks for binary classification using the logistic loss. We show that gradient descent drives the training loss to…
Failures of model-dependent generalization bounds for least-norm interpolation
Peter L. Bartlett, Philip M. Long
We consider bounds on the generalization performance of the least-norm linear regressor, in the over-parameterized regime where it can interpolate the data. We describe a sense in…
Optimal Robust Linear Regression in Nearly Linear Time
Yeshwanth Cherapanamjeri, Efe Aras, Nilesh Tripuraneni +3
We study the problem of high-dimensional robust linear regression where a learner is given access to samples from the generative model (with $X…
On Linear Stochastic Approximation: Fine-grained Polyak-Ruppert and Non-Asymptotic Concentration
Wenlong Mou, Chris Junchi Li, Martin J. Wainwright +2
We undertake a precise study of the asymptotic and non-asymptotic properties of stochastic approximation procedures with Polyak-Ruppert averaging for solving a linear system $\bar{…
Sampling for Bayesian Mixture Models: MCMC with Polynomial-Time Mixing
Wenlong Mou, Nhat Ho, Martin J. Wainwright +2
We study the problem of sampling from the power posterior distribution in Bayesian Gaussian mixture models, a robust version of the classical posterior. This power posterior is kno…