Fantastic Generalization Measures and Where to Find Them
arXiv:1912.02178
Abstract
Generalization of deep networks has been of great interest in recent years, resulting in a number of theoretically and empirically motivated complexity measures. However, most papers proposing such measures study only a small set of models, leaving open the question of whether the conclusion drawn from those experiments would remain valid in other settings. We present the first large scale study of generalization in deep networks. We investigate more then 40 complexity measures taken from both theoretical bounds and empirical studies. We train over 10,000 convolutional networks by systematically varying commonly used hyperparameters. Hoping to uncover potentially causal relationships between each measure and generalization, we analyze carefully controlled experiments and show surprising failures of some measures as well as promising measures for further research.
References in corpus (6)
- On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima
- Regularizing and Optimizing LSTM Language Models
- Regularizing Neural Networks by Penalizing Confident Output Distributions
- Path-SGD: Path-Normalized Optimization in Deep Neural Networks
- Generalization in Deep Networks: The Role of Distance from Initialization
- Deterministic PAC-Bayesian generalization bounds for deep networks via generalizing noise-resilience
Cited by in corpus (15)
- Generalization bounds for deep learning
- Wide flat minima and optimal generalization in classifying high-dimensional Gaussian mixtures
- Traces of Class/Cross-Class Structure Pervade Deep Learning Spectra
- Enhancing Transformers without Self-supervised Learning: A Loss Landscape Perspective in Sequential Recommendation
- A Theoretical Analysis of Fine-tuning with Linear Teachers
- Flatness is a False Friend
- Implicit bias of deep linear networks in the large learning rate phase
- Towards a practical measure of interference for reinforcement learning
- Compressing Heavy-Tailed Weight Matrices for Non-Vacuous Generalization Bounds
- Practical Assessment of Generalization Performance Robustness for Deep Networks via Contrastive Examples
- Generalized Resubstitution for Classification Error Estimation
- Phases of learning dynamics in artificial neural networks: with or without mislabeled data
- Ridgeless Interpolation with Shallow ReLU Networks in is Nearest Neighbor Curvature Extrapolation and Provably Generalizes on Lipschitz Functions
- Exact Stochastic Second Order Deep Learning
- Minimum sharpness: Scale-invariant parameter-robustness of neural networks