Towards Understanding Generalization of Deep Learning: Perspective of Loss Landscapes
arXiv:1706.10239
Abstract
It is widely observed that deep learning models with learned parameters generalize well, even with much more model parameters than the number of training samples. We systematically investigate the underlying reasons why deep neural networks often generalize well, and reveal the difference between the minima (with the same training error) that generalize well and those they don't. We show that it is the characteristics the landscape of the loss function that explains the good generalization capability. For the landscape of loss function for deep networks, the volume of basin of attraction of good minima dominates over that of poor minima, which guarantees optimization methods with random initialization to converge to good minima. We theoretically justify our findings through analyzing 2-layer neural networks; and show that the low-complexity solutions have a small norm of Hessian matrix with respect to model parameters. For deeper networks, extensive numerical evidence helps to support our arguments.
References in corpus (1)
Cited by in corpus (41)
- Shortcut Learning in Deep Neural Networks
- Optimization for deep learning: theory and algorithms
- Overfitting Mechanism and Avoidance in Deep Neural Networks
- Review: Deep Learning in Electron Microscopy
- DeepLesionBrain: Towards a broader deep-learning generalization for multiple sclerosis lesion segmentation
- Universal Effectiveness of High-Depth Circuits in Variational Eigenproblems
- Asymmetric Valleys: Beyond Sharp and Flat Local Minima
- Bridging Mode Connectivity in Loss Landscapes and Adversarial Robustness
- Explicitizing an Implicit Bias of the Frequency Principle in Two-layer Neural Networks
- On the different regimes of Stochastic Gradient Descent
- A Diffusion Theory For Deep Learning Dynamics: Stochastic Gradient Descent Exponentially Favors Flat Minima
- Lexicographic and Depth-Sensitive Margins in Homogeneous and Non-Homogeneous Deep Models
- Luck Matters: Understanding Training Dynamics of Deep ReLU Networks
- Generalization bounds for deep learning
- Which Shortcut Cues Will DNNs Choose? A Study from the Parameter-Space Perspective
- Laziness, Barren Plateau, and Noise in Machine Learning
- Learning Rates as a Function of Batch Size: A Random Matrix Theory Approach to Neural Network Training
- Is SGD a Bayesian sampler? Well, almost
- Emergent Computations in Trained Artificial Neural Networks and Real Brains
- Stochastic Training is Not Necessary for Generalization
- The asymptotic spectrum of the Hessian of DNN throughout training
- Deep learning in bioinformatics: introduction, application, and perspective in big data era
- Understanding Deep Learning Generalization by Maximum Entropy
- Analytic Characterization of the Hessian in Shallow ReLU Models: A Tale of Symmetry
- Flatness is a False Friend
- Understanding over-parameterized deep networks by geometrization
- Why flatness does and does not correlate with generalization for deep neural networks
- Train Flat, Then Compress: Sharpness-Aware Minimization Learns More Compressible Models
- Deforming the Loss Surface to Affect the Behaviour of the Optimizer
- Deep network as memory space: complexity, generalization, disentangled representation and interpretability
- Embedding Principle in Depth for the Loss Landscape Analysis of Deep Neural Networks
- AdaL: Adaptive Gradient Transformation Contributes to Convergences and Generalizations
- Convolutional Gated MLP: Combining Convolutions & gMLP
- Robot Gaining Accurate Pouring Skills through Self-Supervised Learning and Generalization
- Analytic Study of Families of Spurious Minima in Two-Layer ReLU Neural Networks: A Tale of Symmetry II
- Combating Unknown Bias with Effective Bias-Conflicting Scoring and Gradient Alignment
- Deforming the Loss Surface
- Embedding Principle: a hierarchical structure of loss landscape of deep neural networks
- Good Classifiers are Abundant in the Interpolating Regime
- Minimum sharpness: Scale-invariant parameter-robustness of neural networks
- On Large Batch Training and Sharp Minima: A Fokker-Planck Perspective