A Farewell to the Bias-Variance Tradeoff? An Overview of the Theory of Overparameterized Machine Learning
arXiv:2109.02355
Abstract
The last decade of progress in machine learning (ML), especially the deep learning era, has raised a number of scientific questions that challenge the longstanding dogma of the field. One of the most important riddles was the good empirical generalization of overparameterized models. Overparameterized models are highly complex with respect to the size of the training dataset, which enables them to perfectly fit (i.e., interpolate) even noisy training data. Such interpolation of noisy data is traditionally associated with detrimental overfitting, and yet a wide range of interpolating models -- from simple linear models to deep neural networks -- have been observed to generalize remarkably well on fresh test data. Indeed, the discovery of the double descent phenomenon has revealed that highly overparameterized models can improve over the best underparameterized model in test performance. Understanding learning in this overparameterized regime required new theory and foundational empirical studies, even for the simplest case of the linear model. The underpinnings of this understanding have been laid in foundational analyses of overparameterized linear regression and related statistical learning tasks, mostly published between 2018 and 2022, which resulted in precise analytic characterizations of double descent. This paper provides an overview of the theory of overparameterized ML (henceforth abbreviated as TOPML) by focusing on explaining the most foundational findings through a statistical signal processing perspective. We emphasize the unique aspects that define the TOPML research area as a subfield of modern ML theory and outline interesting open frontiers that remain.
References in corpus (27)
- Benign overfitting in ridge regression
- Evaluation of Neural Architectures Trained with Square Loss vs Cross-Entropy in Classification Tasks
- Optimal Regularization Can Mitigate Double Descent
- Classification vs regression in overparameterized regimes: Does the loss function matter?
- Understanding overfitting peaks in generalization error: Analytical risk curves for and penalized interpolation
- Generalization error in high-dimensional perceptrons: Approaching Bayes error with convex optimization
- Minimizing The Misclassification Error Rate Using a Surrogate Convex Loss
- NeurIPS 2020 Competition: Predicting Generalization in Deep Learning
- A Precise Performance Analysis of Learning with Random Features
- Probing transfer learning with a model of synthetic correlated datasets
- Generalization error of random features and kernel methods: hypercontractivity and kernel matrix concentration
- Interpolating Classifiers Make Few Mistakes
- Failures of model-dependent generalization bounds for least-norm interpolation
- Phase Transitions in Transfer Learning for High-Dimensional Perceptrons
- Benign Overfitting in Multiclass Classification: All Roads Lead to Interpolation
- Risk Bounds for Over-parameterized Maximum Margin Classification on Sub-Gaussian Mixtures
- Regularization in High-Dimensional Regression and Classification via Random Matrix Theory
- On the computational and statistical complexity of over-parameterized matrix sensing
- Uniform Convergence of Interpolators: Gaussian Width, Norm Bounds, and Benign Overfitting
- Transfer Learning for Linear Regression: a Statistical Test of Gain
- On Uniform Convergence and Low-Norm Interpolation Learning
- Support vector machines and linear regression coincide with very high-dimensional features
- Recovery and Generalization in Over-Realized Dictionary Learning
- Provable Benefits of Overparameterization in Model Compression: From Double Descent to Pruning Neural Networks
- For interpolating kernel machines, minimizing the norm of the ERM solution minimizes stability
- Double Descent and Other Interpolation Phenomena in GANs
- Distribution of Classification Margins: Are All Data Equal?