Computing Nonvacuous Generalization Bounds for Deep (Stochastic) Neural Networks with Many More Parameters than Training Data
arXiv:1703.11008
Abstract
One of the defining properties of deep learning is that models are chosen to have many more parameters than available training data. In light of this capacity for overfitting, it is remarkable that simple algorithms like SGD reliably return solutions with low test error. One roadblock to explaining these phenomena in terms of implicit regularization, structural properties of the solution, and/or easiness of the data is that many learning bounds are quantitatively vacuous when applied to networks learned by SGD in this "deep learning" regime. Logically, in order to explain generalization, we need nonvacuous bounds. We return to an idea by Langford and Caruana (2001), who used PAC-Bayes bounds to compute nonvacuous numerical bounds on generalization error for stochastic two-layer two-hidden-unit neural networks via a sensitivity analysis. By optimizing the PAC-Bayes bound directly, we are able to extend their approach and obtain nonvacuous generalization bounds for deep stochastic neural network classifiers with millions of parameters trained on only tens of thousands of examples. We connect our findings to recent and old work on flat minima and MDL-based explanations of generalization.
14 pages, 1 table, 2 figures. Corresponds with UAI camera ready and supplement. Includes additional references and related experiments
References in corpus (2)
Cited by in corpus (145)
- Bayesian Deep Convolutional Encoder-Decoder Networks for Surrogate Modeling and Uncertainty Quantification
- Frequency Principle: Fourier Analysis Sheds Light on Deep Neural Networks
- Deep Learning Scaling is Predictable, Empirically
- A deep-learning-based surrogate model for data assimilation in dynamic subsurface flow problems
- Certifying Some Distributional Robustness with Principled Adversarial Training
- Exploring Generalization in Deep Learning
- Neural Lander: Stable Drone Landing Control using Learned Dynamics
- Sensitivity and Generalization in Neural Networks: an Empirical Study
- Fine-Grained Analysis of Optimization and Generalization for Overparameterized Two-Layer Neural Networks
- On the importance of single directions for generalization
- Adversarial Weight Perturbation Helps Robust Generalization
- Towards Understanding the Role of Over-Parametrization in Generalization of Neural Networks
- Bayesian Deep Learning and a Probabilistic Perspective of Generalization
- Fantastic Generalization Measures and Where to Find Them
- Learning ReLU Networks on Linearly Separable Data: Algorithm, Optimality, and Generalization
- Neural Quantum States of frustrated magnets: generalization and sign structure
- A Theoretical Analysis of Deep Q-Learning
- Scalable agent alignment via reward modeling: a research direction
- Sharpness-Aware Minimization for Efficiently Improving Generalization
- The no-free-lunch theorems of supervised learning
- Non-Vacuous Generalization Bounds at the ImageNet Scale: A PAC-Bayesian Compression Approach
- The Pitfalls of Simplicity Bias in Neural Networks
- Stronger generalization bounds for deep nets via a compression approach
- Implicit Regularization in Deep Learning
- Fast Convergence of Natural Gradient Descent for Overparameterized Neural Networks
- An analytic theory of generalization dynamics and transfer learning in deep linear networks
- Identity Crisis: Memorization and Generalization under Extreme Overparameterization
- Classification vs regression in overparameterized regimes: Does the loss function matter?
- High Frequency Component Helps Explain the Generalization of Convolutional Neural Networks
- Reliability and Interpretability in Science and Deep Learning
- Generalization Guarantees for Neural Networks via Harnessing the Low-rank Structure of the Jacobian
- Invariant Causal Prediction for Block MDPs
- Identifying Generalization Properties in Neural Networks
- Tighter risk certificates for neural networks
- Uncertainty Quantification and Deep Ensembles
- Explicitizing an Implicit Bias of the Frequency Principle in Two-layer Neural Networks
- The Impact of Neural Network Overparameterization on Gradient Confusion and Stochastic Gradient Descent
- Analysis of Generalizability of Deep Neural Networks Based on the Complexity of Decision Boundary
- Improved Sample Complexities for Deep Networks and Robust Classification via an All-Layer Margin
- The Implicit and Explicit Regularization Effects of Dropout
- A Diffusion Theory For Deep Learning Dynamics: Stochastic Gradient Descent Exponentially Favors Flat Minima
- Efficient Sharpness-aware Minimization for Improved Training of Neural Networks
- NeurIPS 2020 Competition: Predicting Generalization in Deep Learning
- A Scale Invariant Flatness Measure for Deep Network Minima
- On the role of data in PAC-Bayes bounds
- Assessing Generalization of SGD via Disagreement
- Algorithmic Regularization in Over-parameterized Matrix Sensing and Neural Networks with Quadratic Activations
- On the Benefits of Invariance in Neural Networks
- Observational Overfitting in Reinforcement Learning
- The Deep Bootstrap Framework: Good Online Learners are Good Offline Generalizers
- Representation Based Complexity Measures for Predicting Generalization in Deep Learning
- Learning under Model Misspecification: Applications to Variational and Ensemble methods
- PAC-Bayes with Backprop
- Quantifying the generalization error in deep learning in terms of data distribution and neural network smoothness
- Generalization bounds via distillation
- Distributional Generalization: A New Kind of Generalization
- Minnorm training: an algorithm for training over-parameterized deep neural networks
- A PAC-Bayesian Approach to Generalization Bounds for Graph Neural Networks
- The Surprising Simplicity of the Early-Time Learning Dynamics of Neural Networks
- Benign Overfitting in Multiclass Classification: All Roads Lead to Interpolation
- Beyond Linearization: On Quadratic and Higher-Order Approximation of Wide Neural Networks
- Generalization Bounds for Neural Belief Propagation Decoders
- DNN or k-NN: That is the Generalize vs. Memorize Question
- Heavy Tails in SGD and Compressibility of Overparametrized Neural Networks
- Scalable Marginal Likelihood Estimation for Model Selection in Deep Learning
- Understanding Generalization in Deep Learning via Tensor Methods
- The Training Process of Many Deep Networks Explores the Same Low-Dimensional Manifold
- How noise affects the Hessian spectrum in overparameterized neural networks
- Incorporating Unlabeled Data into Distributionally Robust Learning
- Generalization Guarantees for Imitation Learning
- Wide flat minima and optimal generalization in classifying high-dimensional Gaussian mixtures
- A Farewell to the Bias-Variance Tradeoff? An Overview of the Theory of Overparameterized Machine Learning
- Robustness to Pruning Predicts Generalization in Deep Neural Networks
- A Bayesian Perspective on Training Speed and Model Selection
- Towards Theoretical Understanding of Large Batch Training in Stochastic Gradient Descent
- Sharpness-aware Quantization for Deep Neural Networks
- SALR: Sharpness-aware Learning Rate Scheduler for Improved Generalization
- Gradient-Free Learning Based on the Kernel and the Range Space
- Stationary Points of Shallow Neural Networks with Quadratic Activation Function
- Analytic Network Learning
- For self-supervised learning, Rationality implies generalization, provably
- The Bayesian Learning Rule
- Robustness to Augmentations as a Generalization metric
- Implicit Rugosity Regularization via Data Augmentation
- Deep ReLU Networks Preserve Expected Length
- PAC-Bayes, MAC-Bayes and Conditional Mutual Information: Fast rate bounds that handle general VC classes
- On improving deep learning generalization with adaptive sparse connectivity
- Improving Generalization by Controlling Label-Noise Information in Neural Network Weights
- Neural Complexity Measures
- A Theoretical Analysis of Fine-tuning with Linear Teachers
- PAC-Bayes Analysis Beyond the Usual Bounds
- Generalization Guarantees for Neural Architecture Search with Train-Validation Split
- Capacity Control of ReLU Neural Networks by Basis-path Norm
- PAC-Bayes Bounds for Meta-learning with Data-Dependent Prior
- Data-Dependent Coresets for Compressing Neural Networks with Applications to Generalization Bounds
- Towards Understanding the Generalization Bias of Two Layer Convolutional Linear Classifiers with Gradient Descent
- Sample Complexity Bounds for Recurrent Neural Networks with Application to Combinatorial Graph Problems
- PAC-Bayesian Contrastive Unsupervised Representation Learning
- Still no free lunches: the price to pay for tighter PAC-Bayes bounds
- Exact Gap between Generalization Error and Uniform Convergence in Random Feature Models
- Dichotomize and Generalize: PAC-Bayesian Binary Activated Deep Neural Networks
- A Deep Conditioning Treatment of Neural Networks
- Information-Theoretic Local Minima Characterization and Regularization
- PAC-Bayes Generalisation Bounds for Dynamical Systems Including Stable RNNs
- Fast Convergence for Langevin Diffusion with Manifold Structure
- A Dynamical View on Optimization Algorithms of Overparameterized Neural Networks
- Hyperplane Arrangements of Trained ConvNets Are Biased
- Self-Regularity of Non-Negative Output Weights for Overparameterized Two-Layer Neural Networks
- Noise and Fluctuation of Finite Learning Rate Stochastic Gradient Descent
- Learning Partially Known Stochastic Dynamics with Empirical PAC Bayes
- CAP: Co-Adversarial Perturbation on Weights and Features for Improving Generalization of Graph Neural Networks
- An Optimization and Generalization Analysis for Max-Pooling Networks
- Probably Approximately Correct Vision-Based Planning using Motion Primitives
- Improved generalization by noise enhancement
- Extreme Memorization via Scale of Initialization
- PAC-Bayesian Generalization Bounds for MultiLayer Perceptrons
- A Free-Energy Principle for Representation Learning
- PAC-Bayes Analysis of Sentence Representation
- RATT: Leveraging Unlabeled Data to Guarantee Generalization
- On Predicting Generalization using GANs
- Dissecting Non-Vacuous Generalization Bounds based on the Mean-Field Approximation
- Chaining Meets Chain Rule: Multilevel Entropic Regularization and Training of Neural Nets
- Unifying Variational Inference and PAC-Bayes for Supervised Learning that Scales
- The role of invariance in spectral complexity-based generalization bounds
- Generalisation in fully-connected neural networks for time series forecasting
- Nonlinear Collaborative Scheme for Deep Neural Networks
- Loss Landscape Dependent Self-Adjusting Learning Rates in Decentralized Stochastic Gradient Descent
- Semantically Robust Unpaired Image Translation for Data with Unmatched Semantics Statistics
- Margin Maximization as Lossless Maximal Compression
- On the generalization of bayesian deep nets for multi-class classification
- Uniform Generalization Bounds for Overparameterized Neural Networks
- Inherent Noise in Gradient Based Methods
- Risk Bounds for Learning via Hilbert Coresets
- Maximum Multiscale Entropy and Neural Network Regularization
- Using Human Psychophysics to Evaluate Generalization in Scene Text Recognition Models
- Relative Flatness and Generalization
- Notes on Deep Learning Theory
- Learning Provably Robust Motion Planners Using Funnel Libraries
- A Revision of Neural Tangent Kernel-based Approaches for Neural Networks
- Optimizing Information-theoretical Generalization Bounds via Anisotropic Noise in SGLD
- Characterizing Inter-Layer Functional Mappings of Deep Learning Models
- VC dimension of partially quantized neural networks in the overparametrized regime
- MSR-DARTS: Minimum Stable Rank of Differentiable Architecture Search
- Think Global, Act Local: Relating DNN generalisation and node-level SNR
- Towards Understanding Generalization via Decomposing Excess Risk Dynamics