A Modern Take on the Bias-Variance Tradeoff in Neural Networks
arXiv:1810.08591
Abstract
The bias-variance tradeoff tells us that as model complexity increases, bias falls and variances increases, leading to a U-shaped test error curve. However, recent empirical results with over-parameterized neural networks are marked by a striking absence of the classic U-shaped test error curve: test error keeps decreasing in wider networks. This suggests that there might not be a bias-variance tradeoff in neural networks with respect to network width, unlike was originally claimed by, e.g., Geman et al. (1992). Motivated by the shaky evidence used to support this claim in neural networks, we measure bias and variance in the modern setting. We find that both bias and variance can decrease as the number of parameters grows. To better understand this, we introduce a new decomposition of the variance to disentangle the effects of optimization and data sampling. We also provide theoretical analysis in a simplified setting that is consistent with our empirical findings.
References in corpus (7)
- Reconciling modern machine learning practice and the bias-variance trade-off
- Gradient Descent Provably Optimizes Over-parameterized Neural Networks
- A Closer Look at Memorization in Deep Networks
- Scaling description of generalization with number of parameters in deep learning
- Gradient Descent Happens in a Tiny Subspace
- Implicit Regularization in Deep Learning
- Theory of Deep Learning IIb: Optimization Properties of SGD
Cited by in corpus (31)
- Reconciling modern machine learning practice and the bias-variance trade-off
- A Survey of End-to-End Driving: Architectures and Training Methods
- Scaling description of generalization with number of parameters in deep learning
- Generalisation error in learning with random features and the hidden manifold model
- Double Trouble in Double Descent : Bias and Variance(s) in the Lazy Regime
- Optimal Regularization Can Mitigate Double Descent
- Rethinking Soft Labels for Knowledge Distillation: A Bias-Variance Tradeoff Perspective
- Developments and Further Applications of Ephemeral Data Derived Potentials
- More Data Can Hurt for Linear Regression: Sample-wise Double Descent
- Is deep learning necessary for simple classification tasks?
- The Impact of Neural Network Overparameterization on Gradient Confusion and Stochastic Gradient Descent
- Triple descent and the two kinds of overfitting: Where & why do they appear?
- Multiple Descent: Design Your Own Generalization Curve
- From Stars to Subgraphs: Uplifting Any GNN with Local Structure Awareness
- What causes the test error? Going beyond bias-variance via ANOVA
- On Power Laws in Deep Ensembles
- Random Hypervolume Scalarizations for Provable Multi-Objective Black Box Optimization
- Distributional Generalization: A New Kind of Generalization
- Generalization bounds for deep learning
- Implicit Regularization of Random Feature Models
- Taxonomizing local versus global structure in neural network loss landscapes
- Finding the Needle in the Haystack with Convolutions: on the benefits of architectural bias
- On Sparsity in Overparametrised Shallow ReLU Networks
- Implicit Regularization via Neural Feature Alignment
- Lipschitz Bounds and Provably Robust Training by Laplacian Smoothing
- Perspective: A Phase Diagram for Deep Learning unifying Jamming, Feature Learning and Lazy Training
- On the Universality of the Double Descent Peak in Ridgeless Regression
- Vulnerability Under Adversarial Machine Learning: Bias or Variance?
- About Explicit Variance Minimization: Training Neural Networks for Medical Imaging With Limited Data Annotations
- Mitigating deep double descent by concatenating inputs
- Wavefield solutions from machine learned functions