Eigenvalues of the Hessian in Deep Learning: Singularity and Beyond
arXiv:1611.07476
Abstract
We look at the eigenvalues of the Hessian of a loss function before and after training. The eigenvalue distribution is seen to be composed of two parts, the bulk which is concentrated around zero, and the edges which are scattered away from zero. We present empirical evidence for the bulk indicating how over-parametrized the system is, and for the edges that depend on the input data.
ICLR submission, 2016 - updated to match the openreview.net version
References in corpus (3)
Cited by in corpus (70)
- Prevalence of Neural Collapse during the terminal phase of deep learning training
- Understanding and mitigating gradient pathologies in physics-informed neural networks
- Empirical Analysis of the Hessian of Over-Parametrized Neural Networks
- Optimization for deep learning: theory and algorithms
- Towards Understanding Generalization of Deep Learning: Perspective of Loss Landscapes
- Gradient Descent Happens in a Tiny Subspace
- Perspective: Energy Landscapes for Machine Learning
- Universal discriminative quantum neural networks
- Universal Effectiveness of High-Depth Circuits in Variational Eigenproblems
- Asymptotics of Wide Networks from Feynman Diagrams
- A Closer Look at the Optimization Landscapes of Generative Adversarial Networks
- PyHessian: Neural Networks Through the Lens of the Hessian
- The Full Spectrum of Deepnet Hessians at Scale: Dynamics with SGD Training and Sample Size
- Rethinking Parameter Counting in Deep Models: Effective Dimensionality Revisited
- Measurements of Three-Level Hierarchical Structure in the Outliers in the Spectrum of Deepnet Hessians
- Local Saddle Point Optimization: A Curvature Exploitation Approach
- A literature survey of matrix methods for data science
- Negative eigenvalues of the Hessian in deep neural networks
- Recent advances in deep learning theory
- On the Heavy-Tailed Theory of Stochastic Gradient Descent for Deep Neural Networks
- Deep Learning Through the Lens of Example Difficulty
- Emergent properties of the local geometry of neural loss landscapes
- Weight-space symmetry in deep networks gives rise to permutation saddles, connected by equal-loss valleys across the loss landscape
- Hessian-based toolbox for reliable and interpretable machine learning in physics
- Large Scale Structure of Neural Network Loss Landscapes
- Breaking Reversibility Accelerates Langevin Dynamics for Global Non-Convex Optimization
- Which Algorithmic Choices Matter at Which Batch Sizes? Insights From a Noisy Quadratic Model
- Luck Matters: Understanding Training Dynamics of Deep ReLU Networks
- Directional Pruning of Deep Neural Networks
- Dissecting Hessian: Understanding Common Structure of Hessian in Neural Networks
- Inexact Newton Methods for Stochastic Nonconvex Optimization with Applications to Neural Network Training
- Newton-type Methods for Minimax Optimization
- First Exit Time Analysis of Stochastic Gradient Descent Under Heavy-Tailed Gradient Noise
- How noise affects the Hessian spectrum in overparameterized neural networks
- Learning Rates as a Function of Batch Size: A Random Matrix Theory Approach to Neural Network Training
- Traces of Class/Cross-Class Structure Pervade Deep Learning Spectra
- Taxonomizing local versus global structure in neural network loss landscapes
- On the interplay between noise and curvature and its effect on optimization and generalization
- Hessian Eigenspectra of More Realistic Nonlinear Models
- Beyond Random Matrix Theory for Deep Networks
- GraVAC: Adaptive Compression for Communication-Efficient Distributed DL Training
- Accelerating Distributed ML Training via Selective Synchronization
- Hessian based analysis of SGD for Deep Nets: Dynamics and Generalization
- Analytic Insights into Structure and Rank of Neural Network Hessian Maps
- Convergence rates of Gibbs measures with degenerate minimum
- De-randomized PAC-Bayes Margin Bounds: Applications to Non-convex and Non-smooth Predictors
- SGD in the Large: Average-case Analysis, Asymptotics, and Stepsize Criticality
- Revisiting the Fragility of Influence Functions
- The Limiting Dynamics of SGD: Modified Loss, Phase Space Oscillations, and Anomalous Diffusion
- Learning with Gradient Descent and Weakly Convex Losses
- Analytic Characterization of the Hessian in Shallow ReLU Models: A Tale of Symmetry
- Neumann Optimizer: A Practical Optimization Algorithm for Deep Neural Networks
- Cockpit: A Practical Debugging Tool for the Training of Deep Neural Networks
- Flatness is a False Friend
- Why flatness does and does not correlate with generalization for deep neural networks
- Deep Curvature Suite
- Low Rank Saddle Free Newton: A Scalable Method for Stochastic Nonconvex Optimization
- An analytic theory of shallow networks dynamics for hinge loss classification
- Eigencurve: Optimal Learning Rate Schedule for SGD on Quadratic Objectives with Skewed Hessian Spectrums
- On the Convex Behavior of Deep Neural Networks in Relation to the Layers' Width
- How many degrees of freedom do we need to train deep networks: a loss landscape perspective
- ViViT: Curvature access through the generalized Gauss-Newton's low-rank structure
- Appearance of Random Matrix Theory in Deep Learning
- Mean Field Theory of Activation Functions in Deep Neural Networks
- Mode connectivity in the QCBM loss landscape
- Riemannian Laplace approximations for Bayesian neural networks
- Analytic Study of Families of Spurious Minima in Two-Layer ReLU Neural Networks: A Tale of Symmetry II
- Run-and-Inspect Method for Nonconvex Optimization and Global Optimality Bounds for R-Local Minimizers
- Dynamics of Stochastic Momentum Methods on Large-scale, Quadratic Models
- Geometry Perspective Of Estimating Learning Capability Of Neural Networks