What causes the test error? Going beyond bias-variance via ANOVA
arXiv:2010.05170
Abstract
Modern machine learning methods are often overparametrized, allowing adaptation to the data at a fine level. This can seem puzzling; in the worst case, such models do not need to generalize. This puzzle inspired a great amount of work, arguing when overparametrization reduces test error, in a phenomenon called "double descent". Recent work aimed to understand in greater depth why overparametrization is helpful for generalization. This leads to discovering the unimodality of variance as a function of the level of parametrization, and to decomposing the variance into that arising from label noise, initialization, and randomness in the training data to understand the sources of the error. In this work we develop a deeper understanding of this area. Specifically, we propose using the analysis of variance (ANOVA) to decompose the variance in the test error in a symmetric way, for studying the generalization performance of certain two-layer linear and non-linear networks. The advantage of the analysis of variance is that it reveals the effects of initialization, label noise, and training data more clearly than prior approaches. Moreover, we also study the monotonicity and unimodality of the variance components. While prior work studied the unimodality of the overall variance, we study the properties of each term in variance decomposition. One key insight is that in typical settings, the interaction between training samples and initialization can dominate the variance; surprisingly being larger than their marginal effect. Also, we characterize "phase transitions" where the variance changes from unimodal to monotone. On a technical level, we leverage advanced deterministic equivalent techniques for Haar random matrices, that -- to our knowledge -- have not yet been used in the area. We also verify our results in numerical simulations and on empirical data examples.
References in corpus (21)
- Wide Residual Networks
- Reconciling modern machine learning practice and the bias-variance trade-off
- Benign Overfitting in Linear Regression
- The generalization error of random features regression: Precise asymptotics and double descent curve
- Scaling description of generalization with number of parameters in deep learning
- Overfitting or perfect fitting? Risk bounds for classification and regression rules that interpolate
- Memorizing without overfitting: Bias, variance, and interpolation in over-parameterized models
- A Modern Take on the Bias-Variance Tradeoff in Neural Networks
- More Data Can Hurt for Linear Regression: Sample-wise Double Descent
- A Large Dimensional Analysis of Least Squares Support Vector Machines
- Random Beamforming over Quasi-Static and Fading Channels: A Deterministic Equivalent Approach
- Provable Benefit of Orthogonal Initialization in Optimizing Deep Linear Networks
- The Neural Tangent Kernel in High Dimensions: Triple Descent and a Multi-Scale Theory of Generalization
- Exact expressions for double descent and implicit regularization via surrogate random design
- Multiple Descent: Design Your Own Generalization Curve
- Deep Isometric Learning for Visual Recognition
- Understanding Double Descent Requires a Fine-Grained Bias-Variance Decomposition
- Distributed linear regression by averaging
- A Random Matrix Perspective on Mixtures of Nonlinearities for Deep Learning
- On the Peaking Phenomenon of the Lasso in Model Selection
- Provable More Data Hurt in High Dimensional Least Squares Estimator
Cited by in corpus (10)
- The Modern Mathematics of Deep Learning
- Memorizing without overfitting: Bias, variance, and interpolation in over-parameterized models
- Random Features for Kernel Approximation: A Survey on Algorithms, Theory, and Beyond
- Understanding Generalization in Adversarial Training via the Bias-Variance Decomposition
- Dimensionality reduction, regularization, and generalization in overparameterized regressions
- On the Double Descent of Random Features Models Trained with SGD
- The Geometry of Over-parameterized Regression and Adversarial Perturbations
- Deformed semicircle law and concentration of nonlinear random matrices for ultra-wide neural networks
- Covariate Shift in High-Dimensional Random Feature Regression
- Harmless interpolation in regression and classification with structured features