Sharp Minima Can Generalize For Deep Nets
arXiv:1703.04933
Abstract
Despite their overwhelming capacity to overfit, deep learning architectures tend to generalize relatively well to unseen data, allowing them to be deployed in practice. However, explaining why this is the case is still an open area of research. One standing hypothesis that is gaining popularity, e.g. Hochreiter & Schmidhuber (1997); Keskar et al. (2017), is that the flatness of minima of the loss function found by stochastic gradient based methods results in good generalization. This paper argues that most notions of flatness are problematic for deep models and can not be directly applied to explain generalization. Specifically, when focusing on deep networks with rectifier units, we can exploit the particular geometry of parameter space induced by the inherent symmetries that these architectures exhibit to build equivalent models corresponding to arbitrarily sharper minima. Furthermore, if we allow to reparametrize a function, the geometry of its parameters can change drastically without affecting its generalization properties.
8.5 pages of main content, 2.5 of bibliography and 1 page of appendix
References in corpus (14)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Sequence to Sequence Learning with Neural Networks
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation
- Attention-Based Models for Speech Recognition
- Improved Techniques for Training GANs
- Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification
- Density estimation using Real NVP
- The Loss Surfaces of Multilayer Networks
- Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
- Wav2Letter: an End-to-End ConvNet-based Speech Recognition System
- Path-SGD: Path-Normalized Optimization in Deep Neural Networks
- A Convolutional Encoder Model for Neural Machine Translation
- Natural Neural Networks
Cited by in corpus (88)
- Averaging Weights Leads to Wider Optima and Better Generalization
- Don't Use Large Mini-Batches, Use Local SGD
- Fantastic Generalization Measures and Where to Find Them
- Measuring the Effects of Data Parallelism on Neural Network Training
- Optimization for deep learning: theory and algorithms
- Deep learning generalizes because the parameter-function map is biased towards simple functions
- The large learning rate phase of deep learning: the catapult mechanism
- Generalization in Deep Networks: The Role of Distance from Initialization
- Where is the Information in a Deep Neural Network?
- Universal Statistics of Fisher Information in Deep Neural Networks: Mean Field Approach
- High Frequency Component Helps Explain the Generalization of Convolutional Neural Networks
- The Heavy-Tail Phenomenon in SGD
- Explicit Regularisation in Gaussian Noise Injections
- On the Loss Landscape of Adversarial Training: Identifying Challenges and How to Overcome Them
- Implicit Bias of Gradient Descent on Linear Convolutional Networks
- Is it enough to optimize CNN architectures on ImageNet?
- A Diffusion Theory For Deep Learning Dynamics: Stochastic Gradient Descent Exponentially Favors Flat Minima
- Efficient Sharpness-aware Minimization for Improved Training of Neural Networks
- Understanding Generalization through Visualizations
- Recent advances in deep learning theory
- Time Matters in Regularizing Deep Networks: Weight Decay and Data Augmentation Affect Early Learning Dynamics, Matter Little Near Convergence
- ASAM: Adaptive Sharpness-Aware Minimization for Scale-Invariant Learning of Deep Neural Networks
- Emergent properties of the local geometry of neural loss landscapes
- On the Validity of Modeling SGD with Stochastic Differential Equations (SDEs)
- Lipschitzness Is All You Need To Tame Off-policy Generative Adversarial Imitation Learning
- Stochastic Weight Averaging in Parallel: Large-Batch Training that Generalizes Well
- Quantifying the generalization error in deep learning in terms of data distribution and neural network smoothness
- Distributed Learning of Deep Neural Networks using Independent Subnet Training
- Normalized Flat Minima: Exploring Scale Invariant Definition of Flat Minima for Neural Networks using PAC-Bayesian Analysis
- Wide-minima Density Hypothesis and the Explore-Exploit Learning Rate Schedule
- Luck Matters: Understanding Training Dynamics of Deep ReLU Networks
- Generalization in Machine Learning via Analytical Learning Theory
- PAC-Bayes Information Bottleneck
- Reconciling Modern Deep Learning with Traditional Optimization Analyses: The Intrinsic Learning Rate
- Visualized Insights into the Optimization Landscape of Fully Convolutional Networks
- Learning Rates as a Function of Batch Size: A Random Matrix Theory Approach to Neural Network Training
- Learning Rate Annealing Can Provably Help Generalization, Even for Convex Problems
- Information-Theoretic Generalization Bounds for Stochastic Gradient Descent
- Orthogonal Deep Neural Networks
- Taxonomizing local versus global structure in neural network loss landscapes
- Empirical study of extreme overfitting points of neural networks
- Is SGD a Bayesian sampler? Well, almost
- On the interplay between noise and curvature and its effect on optimization and generalization
- SALR: Sharpness-aware Learning Rate Scheduler for Improved Generalization
- StylePredict: Machine Theory of Mind for Human Driver Behavior From Trajectories
- Artificial Neural Variability for Deep Learning: On Overfitting, Noise Memorization, and Catastrophic Forgetting
- Regularizing Neural Networks via Adversarial Model Perturbation
- Hessian Eigenspectra of More Realistic Nonlinear Models
- Adversarial Training Makes Weight Loss Landscape Sharper in Logistic Regression
- Information-Theoretic Local Minima Characterization and Regularization
- The Representation Theory of Neural Networks
- De-randomized PAC-Bayes Margin Bounds: Applications to Non-convex and Non-smooth Predictors
- Towards Better Plasticity-Stability Trade-off in Incremental Learning: A Simple Linear Connector
- Analytic Characterization of the Hessian in Shallow ReLU Models: A Tale of Symmetry
- Why flatness does and does not correlate with generalization for deep neural networks
- Tangent Space Separability in Feedforward Neural Networks
- Flatness is a False Friend
- CAP: Co-Adversarial Perturbation on Weights and Features for Improving Generalization of Graph Neural Networks
- Does the Data Induce Capacity Control in Deep Learning?
- Neighborhood-Aware Neural Architecture Search
- Cockpit: A Practical Debugging Tool for the Training of Deep Neural Networks
- Stochasticity of Deterministic Gradient Descent: Large Learning Rate for Multiscale Objective Function
- Smoothness Analysis of Adversarial Training
- On the Generalization of Models Trained with SGD: Information-Theoretic Bounds and Implications
- What training reveals about neural network complexity
- Optimization Variance: Exploring Generalization Properties of DNNs
- Deforming the Loss Surface to Affect the Behaviour of the Optimizer
- A Reparameterization-Invariant Flatness Measure for Deep Neural Networks
- Adaptive Periodic Averaging: A Practical Approach to Reducing Communication in Distributed Learning
- Mean Field Theory of Activation Functions in Deep Neural Networks
- Explicitly Bayesian Regularizations in Deep Learning
- Unique Properties of Flat Minima in Deep Networks
- On regularization of gradient descent, layer imbalance and flat minima
- Ensemble Feature for Person Re-Identification
- Dissecting Non-Vacuous Generalization Bounds based on the Mean-Field Approximation
- BN-invariant sharpness regularizes the training model to better generalization
- TATi-Thermodynamic Analytics ToolkIt: TensorFlow-based software for posterior sampling in machine learning applications
- Inherent Noise in Gradient Based Methods
- Deforming the Loss Surface
- Training Deep Neural Networks via Branch-and-Bound
- Stochastic Function Norm Regularization of Deep Networks
- Relative Flatness and Generalization
- On Large Batch Training and Sharp Minima: A Fokker-Planck Perspective
- Adaptive Weight Decay for Deep Neural Networks
- Minimum sharpness: Scale-invariant parameter-robustness of neural networks
- Analytic Study of Families of Spurious Minima in Two-Layer ReLU Neural Networks: A Tale of Symmetry II
- Robot Gaining Accurate Pouring Skills through Self-Supervised Learning and Generalization
- Geometry Perspective Of Estimating Learning Capability Of Neural Networks