Triple descent and the two kinds of overfitting: Where & why do they appear?
arXiv:2006.03509 · doi:10.1088/1742-5468/ac3909
Abstract
A recent line of research has highlighted the existence of a "double descent" phenomenon in deep learning, whereby increasing the number of training examples causes the generalization error of neural networks to peak when is of the same order as the number of parameters . In earlier works, a similar phenomenon was shown to exist in simpler models such as linear regression, where the peak instead occurs when is equal to the input dimension . Since both peaks coincide with the interpolation threshold, they are often conflated in the litterature. In this paper, we show that despite their apparent similarity, these two scenarios are inherently different. In fact, both peaks can co-exist when neural networks are applied to noisy regression tasks. The relative size of the peaks is then governed by the degree of nonlinearity of the activation function. Building on recent developments in the analysis of random feature models, we provide a theoretical ground for this sample-wise triple descent. As shown previously, the nonlinear peak at is a true divergence caused by the extreme sensitivity of the output function to both the noise corrupting the labels and the initialization of the random features (or the weights in neural networks). This peak survives in the absence of noise, but can be suppressed by regularization. In contrast, the linear peak at is solely due to overfitting the noise in the labels, and forms earlier during training. We show that this peak is implicitly regularized by the nonlinearity, which is why it only becomes salient at high noise and is weakly affected by explicit regularization. Throughout the paper, we compare analytical results obtained in the random feature model with the outcomes of numerical experiments involving deep neural networks.
References in corpus (15)
- Sequence to Sequence Learning with Neural Networks
- Understanding deep learning requires rethinking generalization
- The generalization error of random features regression: Precise asymptotics and double descent curve
- Bayesian Deep Learning and a Probabilistic Perspective of Generalization
- A Landscape Analysis of Constraint Satisfaction Problems
- Generalisation error in learning with random features and the hidden manifold model
- Double Trouble in Double Descent : Bias and Variance(s) in the Lazy Regime
- Optimal Regularization Can Mitigate Double Descent
- The Gaussian equivalence of generative models for learning with shallow neural networks
- More Data Can Hurt for Linear Regression: Sample-wise Double Descent
- The Curious Case of Adversarially Robust Models: More Data Can Help, Double Descend, or Hurt Generalization
- Generalization error in high-dimensional perceptrons: Approaching Bayes error with convex optimization
- Multiple Descent: Design Your Own Generalization Curve
- A Random Matrix Perspective on Mixtures of Nonlinearities for Deep Learning
- Provable More Data Hurt in High Dimensional Least Squares Estimator
Cited by in corpus (10)
- Memorizing without overfitting: Bias, variance, and interpolation in over-parameterized models
- Optimal Regularization Can Mitigate Double Descent
- Multiple Descent: Design Your Own Generalization Curve
- Dimensionality reduction, regularization, and generalization in overparameterized regressions
- Taxonomizing local versus global structure in neural network loss landscapes
- On the Universality of the Double Descent Peak in Ridgeless Regression
- Trap of Feature Diversity in the Learning of MLPs
- The Geometry of Over-parameterized Regression and Adversarial Perturbations
- Risk-Monotonicity in Statistical Learning
- Out-of-Distribution Generalization in Kernel Regression