Rethinking generalization requires revisiting old ideas: statistical mechanics approaches and complex learning behavior
arXiv:1710.09553
Abstract
We describe an approach to understand the peculiar and counterintuitive generalization properties of deep neural networks. The approach involves going beyond worst-case theoretical capacity control frameworks that have been popular in machine learning in recent years to revisit old ideas in the statistical mechanics of neural networks. Within this approach, we present a prototypical Very Simple Deep Learning (VSDL) model, whose behavior is controlled by two control parameters, one describing an effective amount of data, or load, on the network (that decreases when noise is added to the input), and one with an effective temperature interpretation (that increases when algorithms are early stopped). Using this model, we describe how a very simple application of ideas from the statistical mechanics theory of generalization provides a strong qualitative description of recently-observed empirical results regarding the inability of deep neural networks not to overfit training data, discontinuous learning and sharp transitions in the generalization properties of learning algorithms, etc.
31 pages; added brief discussion of recent papers that use/extend these ideas
References in corpus (14)
- Opening the Black Box of Deep Neural Networks via Information
- The Loss Surfaces of Multilayer Networks
- Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
- Deep Learning is Robust to Massive Label Noise
- Qualitatively characterizing neural network optimization problems
- Spectrally-normalized margin bounds for neural networks
- Measuring the Effects of Data Parallelism on Neural Network Training
- Implicit Self-Regularization in Deep Neural Networks: Evidence from Random Matrix Theory and Implications for Learning
- Universal adversarial perturbations
- On the saddle point problem for non-convex optimization
- On the Computational Inefficiency of Large Batch Sizes for Stochastic Gradient Descent
- A Surprising Linear Relationship Predicts Test Performance in Deep Networks
- Theory IIIb: Generalization in Deep Networks
- A Capacity Scaling Law for Artificial Neural Networks
Cited by in corpus (15)
- Optimal Errors and Phase Transitions in High-Dimensional Generalized Linear Models
- Predicting trends in the quality of state-of-the-art neural networks without access to training or testing data
- Toward Convolutional Blind Denoising of Real Photographs
- Training behavior of deep neural network in frequency domain
- Traditional and Heavy-Tailed Self Regularization in Neural Network Models
- Towards Efficient Training for Neural Network Quantization
- Multiplicative noise and heavy tails in stochastic optimization
- On Random Matrices Arising in Deep Neural Networks. Gaussian Case
- A Random Matrix Analysis of Random Fourier Features: Beyond the Gaussian Kernel, a Precise Phase Transition, and the Corresponding Double Descent
- Invariance of Weight Distributions in Rectified MLPs
- Taxonomizing local versus global structure in neural network loss landscapes
- Parameter Re-Initialization through Cyclical Batch Size Schedules
- A study of CNN capacity applied to Left Venticle Segmentation in Cardiac MRI
- Eigenvalue Distribution of Large Random Matrices Arising in Deep Neural Networks: Orthogonal Case
- Good Classifiers are Abundant in the Interpolating Regime