Towards an Understanding of Benign Overfitting in Neural Networks
arXiv:2106.03212
Abstract
Modern machine learning models often employ a huge number of parameters and are typically optimized to have zero training loss; yet surprisingly, they possess near-optimal prediction performance, contradicting classical learning theory. We examine how these benign overfitting phenomena occur in a two-layer neural network setting where sample covariates are corrupted with noise. We address the high dimensional regime, where the data dimension grows with the number of data points. Our analysis combines an upper bound on the bias with matching upper and lower bounds on the variance of the interpolator (an estimator that interpolates the data). These results indicate that the excess learning risk of the interpolator decays under mild conditions. We further show that it is possible for the two-layer ReLU network interpolator to achieve a near minimax-optimal learning rate, which to our knowledge is the first generalization result for such networks. Finally, our theory predicts that the excess learning risk starts to increase once the number of parameters grows beyond , matching recent empirical findings.
References in corpus (5)
- The Neural Tangent Kernel in High Dimensions: Triple Descent and a Multi-Scale Theory of Generalization
- Multiple Descent: Design Your Own Generalization Curve
- Risk Bounds for Over-parameterized Maximum Margin Classification on Sub-Gaussian Mixtures
- Risk of the Least Squares Minimum Norm Estimator under the Spike Covariance Model
- Benign Overfitting and Noisy Features
Cited by in corpus (4)
- On the Double Descent of Random Features Models Trained with SGD
- Mean-field Analysis of Piecewise Linear Solutions for Wide ReLU Networks
- Harmless interpolation in regression and classification with structured features
- The Interplay Between Implicit Bias and Benign Overfitting in Two-Layer Linear Networks