17 citations · 25 across the 3 of their papers we have counts for
11 papers
When does gradient descent with logistic loss interpolate using deep networks with smoothed ReLU activations?
Niladri S. Chatterji, Philip M. Long, Peter L. Bartlett
We establish conditions under which gradient descent applied to fixed-width deep networks drives the logistic loss to zero, and prove bounds on the rate of convergence. Our analysi…
When does gradient descent with logistic loss find interpolating two-layer networks?
Niladri S. Chatterji, Philip M. Long, Peter L. Bartlett
We study the training of finite-width two-layer smoothed ReLU networks for binary classification using the logistic loss. We show that gradient descent drives the training loss to…
Failures of model-dependent generalization bounds for least-norm interpolation
Peter L. Bartlett, Philip M. Long
We consider bounds on the generalization performance of the least-norm linear regressor, in the over-parameterized regime where it can interpolate the data. We describe a sense in…
On the Global Convergence of Training Deep Linear ResNets
Difan Zou, Philip M. Long, Quanquan Gu
We study the convergence of gradient descent (GD) and stochastic gradient descent (SGD) for training -hidden-layer linear residual networks (ResNets). We prove that for training…
Generalization bounds for deep convolutional neural networks
Philip M. Long, Hanie Sedghi
We prove bounds on the generalization error of convolutional networks. The bounds are in terms of the training loss, the number of parameters, the Lipschitz constant of the loss an…
On the effect of the activation function on the distribution of hidden nodes in a deep network
Philip M. Long, Hanie Sedghi
We analyze the joint probability distribution on the lengths of the vectors of hidden variables in different layers of a fully connected deep network, when the weights and biases a…