Surprises in High-Dimensional Ridgeless Least Squares Interpolation
arXiv:1903.08560
Abstract
Interpolators -- estimators that achieve zero training error -- have attracted growing attention in machine learning, mainly because state-of-the art neural networks appear to be models of this type. In this paper, we study minimum norm ("ridgeless") interpolation in high-dimensional least squares regression. We consider two different models for the feature distribution: a linear model, where the feature vectors are obtained by applying a linear transform to a vector of i.i.d. entries, (with ); and a nonlinear model, where the feature vectors are obtained by passing the input through a random one-layer neural network, (with , a matrix of i.i.d. entries, and an activation function acting componentwise on ). We recover -- in a precise quantitative way -- several phenomena that have been observed in large-scale neural networks and kernel machines, including the "double descent" behavior of the prediction risk, and the potential benefits of overparametrization.
68 pages; 16 figures. This revision contains non-asymptotic version of earlier results, and results for general coefficients
Cited by in corpus (17)
- Scaling description of generalization with number of parameters in deep learning
- The Modern Mathematics of Deep Learning
- Generalisation error in learning with random features and the hidden manifold model
- Memorizing without overfitting: Bias, variance, and interpolation in over-parameterized models
- Learning curves of generic features maps for realistic datasets with a teacher-student model
- Gaussian Universality of Perceptrons with Random Labels
- Boolean learning under noise-perturbations in hardware neural networks
- An Assumption-Free Exact Test For Fixed-Design Linear Models With Exchangeable Errors
- Optimal regularizations for data generation with probabilistic graphical models
- Generalization Error Rates in Kernel Regression: The Crossover from the Noiseless to Noisy Regime
- Dimensionality reduction, regularization, and generalization in overparameterized regressions
- Ridge-type Linear Shrinkage Estimation of the Matrix Mean of High-dimensional Normal Distribution
- On the Generalization Power of Overfitted Two-Layer Neural Tangent Kernel Models
- Ridgeless Regression with Random Features
- The loss landscape of deep linear neural networks: a second-order analysis
- Over-Parametrized Matrix Factorization in the Presence of Spurious Stationary Points
- Dropout Drops Double Descent