Uniform convergence may be unable to explain generalization in deep learning
arXiv:1902.04742
Abstract
Aimed at explaining the surprisingly good generalization behavior of overparameterized deep networks, recent works have developed a variety of generalization bounds for deep learning, all based on the fundamental learning-theoretic technique of uniform convergence. While it is well-known that many of these existing bounds are numerically large, through numerous experiments, we bring to light a more concerning aspect of these bounds: in practice, these bounds can {\em increase} with the training dataset size. Guided by our observations, we then present examples of overparameterized linear classifiers and neural networks trained by gradient descent (GD) where uniform convergence provably cannot "explain generalization" -- even if we take into account the implicit bias of GD {\em to the fullest extent possible}. More precisely, even if we consider only the set of classifiers output by GD, which have test errors less than some small in our settings, we show that applying (two-sided) uniform convergence on this set of classifiers will yield only a vacuous generalization guarantee larger than . Through these findings, we cast doubt on the power of uniform convergence-based generalization bounds to provide a complete picture of why overparameterized deep networks generalize well.
Cited by in corpus (68)
- Stabilizing the Lottery Ticket Hypothesis
- The Modern Mathematics of Deep Learning
- Linear Mode Connectivity and the Lottery Ticket Hypothesis
- What Neural Networks Memorize and Why: Discovering the Long Tail via Influence Estimation
- Self-Adaptive Training: beyond Empirical Risk Minimization
- The Pitfalls of Simplicity Bias in Neural Networks
- Understanding the Failure Modes of Out-of-Distribution Generalization
- Classification vs regression in overparameterized regimes: Does the loss function matter?
- Rethinking Soft Labels for Knowledge Distillation: A Bias-Variance Tradeoff Perspective
- Dataset Inference: Ownership Resolution in Machine Learning
- Finite-sample Analysis of Interpolating Linear Classifiers in the Overparameterized Regime
- Coherent Gradients: An Approach to Understanding Generalization in Gradient Descent-based Optimization
- Assessing Generalization of SGD via Disagreement
- Failures of model-dependent generalization bounds for least-norm interpolation
- What can linearized neural networks actually say about generalization?
- Generalization bounds via distillation
- Distributional Generalization: A New Kind of Generalization
- Toward Better Generalization Bounds with Locally Elastic Stability
- Generalization bounds for deep learning
- Hold me tight! Influence of discriminative features on deep network boundaries
- A Bayesian Perspective on Training Speed and Model Selection
- A Farewell to the Bias-Variance Tradeoff? An Overview of the Theory of Overparameterized Machine Learning
- Uniform Convergence of Interpolators: Gaussian Width, Norm Bounds, and Benign Overfitting
- For self-supervised learning, Rationality implies generalization, provably
- Evaluation of Complexity Measures for Deep Learning Generalization in Medical Image Analysis
- Can Implicit Bias Explain Generalization? Stochastic Convex Optimization as a Case Study
- Self-Adaptive Training: Bridging Supervised and Self-Supervised Learning
- On Uniform Convergence and Low-Norm Interpolation Learning
- Formalizing Generalization and Robustness of Neural Networks to Weight Perturbations
- Train simultaneously, generalize better: Stability of gradient-based minimax learners
- Weak and Strong Gradient Directions: Explaining Memorization, Generalization, and Hardness of Examples at Scale
- Measuring Generalization with Optimal Transport
- Exact Gap between Generalization Error and Uniform Convergence in Random Feature Models
- De-randomized PAC-Bayes Margin Bounds: Applications to Non-convex and Non-smooth Predictors
- Improved Generalization Bounds of Group Invariant / Equivariant Deep Networks via Quotient Feature Spaces
- Making learning more transparent using conformalized performance prediction
- Empirical Risk Minimization in the Interpolating Regime with Application to Neural Network Learning
- In Search of Robust Measures of Generalization
- CoReS: Compatible Representations via Stationarity
- Towards Optimal Problem Dependent Generalization Error Bounds in Statistical Learning Theory
- Learning Theory Can (Sometimes) Explain Generalisation in Graph Neural Networks
- A Learning Theoretic Perspective on Local Explainability
- A PAC-Bayes Analysis of Adversarial Robustness
- Sub-Optimal Local Minima Exist for Neural Networks with Almost All Non-Linear Activations
- A Probabilistic Representation of DNNs: Bridging Mutual Information and Generalization
- RATT: Leveraging Unlabeled Data to Guarantee Generalization
- Making Coherence Out of Nothing At All: Measuring the Evolution of Gradient Alignment
- Functional Regularization for Representation Learning: A Unified Theoretical Perspective
- PAC-Bayesian Generalization Bounds for MultiLayer Perceptrons
- Uniform Convergence, Adversarial Spheres and a Simple Remedy
- Learning Neural Models for Natural Language Processing in the Face of Distributional Shift
- Asymptotic Risk of Overparameterized Likelihood Models: Double Descent Theory for Deep Neural Networks
- Dissecting Non-Vacuous Generalization Bounds based on the Mean-Field Approximation
- The role of invariance in spectral complexity-based generalization bounds
- Realizable Learning is All You Need
- Topologically Densified Distributions
- Risk Bounds for Learning via Hilbert Coresets
- Tangent Space Sensitivity and Distribution of Linear Regions in ReLU Networks
- Generalization by Recognizing Confusion
- Learning the mapping : the cost of finding the needle in a haystack
- Relative Flatness and Generalization
- Good Classifiers are Abundant in the Interpolating Regime
- The Relativity of Induction
- From deep to Shallow: Equivalent Forms of Deep Networks in Reproducing Kernel Krein Space and Indefinite Support Vector Machines
- Covariate Shift in High-Dimensional Random Feature Regression
- Ghosts in Neural Networks: Existence, Structure and Role of Infinite-Dimensional Null Space
- On the Regularization of Autoencoders
- Notes on Deep Learning Theory