activity
20122021
most citedActive and passive learning of linear separators under log-concave distributions

17 citations · 25 across the 3 of their papers we have counts for

collaborators

11 papers

stat.ML2021

When does gradient descent with logistic loss interpolate using deep networks with smoothed ReLU activations?

Niladri S. Chatterji, Philip M. Long, Peter L. Bartlett

We establish conditions under which gradient descent applied to fixed-width deep networks drives the logistic loss to zero, and prove bounds on the rate of convergence. Our analysi…

stat.ML2020

When does gradient descent with logistic loss find interpolating two-layer networks?

Niladri S. Chatterji, Philip M. Long, Peter L. Bartlett

We study the training of finite-width two-layer smoothed ReLU networks for binary classification using the logistic loss. We show that gradient descent drives the training loss to…

stat.ML2020

Failures of model-dependent generalization bounds for least-norm interpolation

Peter L. Bartlett, Philip M. Long

We consider bounds on the generalization performance of the least-norm linear regressor, in the over-parameterized regime where it can interpolate the data. We describe a sense in…

cs.LG20208 cited

On the Global Convergence of Training Deep Linear ResNets

Difan Zou, Philip M. Long, Quanquan Gu

We study the convergence of gradient descent (GD) and stochastic gradient descent (SGD) for training -hidden-layer linear residual networks (ResNets). We prove that for training…

cs.LG2019

Generalization bounds for deep convolutional neural networks

Philip M. Long, Hanie Sedghi

We prove bounds on the generalization error of convolutional networks. The bounds are in terms of the training loss, the number of parameters, the Lipschitz constant of the loss an…

cs.LG2019

On the effect of the activation function on the distribution of hidden nodes in a deep network

Philip M. Long, Hanie Sedghi

We analyze the joint probability distribution on the lengths of the vectors of hidden variables in different layers of a fully connected deep network, when the weights and biases a…