Empirical Analysis of the Hessian of Over-Parametrized Neural Networks
arXiv:1706.04454
Abstract
We study the properties of common loss surfaces through their Hessian matrix. In particular, in the context of deep learning, we empirically show that the spectrum of the Hessian is composed of two parts: (1) the bulk centered near zero, (2) and outliers away from the bulk. We present numerical evidence and mathematical justifications to the following conjectures laid out by Sagun et al. (2016): Fixing data, increasing the number of parameters merely scales the bulk of the spectrum; fixing the dimension and changing the data (for instance adding more clusters or making the data less separable) only affects the outliers. We believe that our observations have striking implications for non-convex optimization in high dimensions. First, the flatness of such landscapes (which can be measured by the singularity of the Hessian) implies that classical notions of basins of attraction may be quite misleading. And that the discussion of wide/narrow basins may be in need of a new perspective around over-parametrization and redundancy that are able to create large connected components at the bottom of the landscape. Second, the dependence of small number of large eigenvalues to the data distribution can be linked to the spectrum of the covariance matrix of gradients of model outputs. With this in mind, we may reevaluate the connections within the data-architecture-algorithm framework of a model, hoping that it would shed light into the geometry of high-dimensional and non-convex spaces in modern applications. In particular, we present a case that links the two observations: small and large batch gradient descent appear to converge to different basins of attraction but we show that they are in fact connected through their flat region and so belong to the same basin.
Minor update for ICLR 2018 Workshop Track presentation
References in corpus (8)
- Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
- Understanding deep learning requires rethinking generalization
- Opening the Black Box of Deep Neural Networks via Information
- On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima
- No More Pesky Learning Rates
- Three Factors Influencing Minima in SGD
- Perspective: Energy Landscapes for Machine Learning
- The Landscape of Empirical Risk for Non-convex Losses
Cited by in corpus (85)
- Prevalence of Neural Collapse during the terminal phase of deep learning training
- Scaling description of generalization with number of parameters in deep learning
- Gradient Descent with Early Stopping is Provably Robust to Label Noise for Overparameterized Neural Networks
- Optimization for deep learning: theory and algorithms
- Deep learning generalizes because the parameter-function map is biased towards simple functions
- Gradient Descent Happens in a Tiny Subspace
- To understand deep learning we need to understand kernel learning
- Towards moderate overparameterization: global convergence guarantees for training shallow neural networks
- Memorizing without overfitting: Bias, variance, and interpolation in over-parameterized models
- A Modern Take on the Bias-Variance Tradeoff in Neural Networks
- Towards Theoretically Understanding Why SGD Generalizes Better Than ADAM in Deep Learning
- Energy-entropy competition and the effectiveness of stochastic gradient descent in machine learning
- The Early Phase of Neural Network Training
- Lipschitz Recurrent Neural Networks
- Accelerating SGD with momentum for over-parameterized learning
- Universal Statistics of Fisher Information in Deep Neural Networks: Mean Field Approach
- Asymptotics of Wide Networks from Feynman Diagrams
- Generalization Guarantees for Neural Networks via Harnessing the Low-rank Structure of the Jacobian
- Explicit Regularisation in Gaussian Noise Injections
- Asymmetric Valleys: Beyond Sharp and Flat Local Minima
- Variational Quantum Classifiers Through the Lens of the Hessian
- La-MAML: Look-ahead Meta Learning for Continual Learning
- The Break-Even Point on Optimization Trajectories of Deep Neural Networks
- The Impact of Neural Network Overparameterization on Gradient Confusion and Stochastic Gradient Descent
- A Closer Look at the Optimization Landscapes of Generative Adversarial Networks
- PyHessian: Neural Networks Through the Lens of the Hessian
- The Full Spectrum of Deepnet Hessians at Scale: Dynamics with SGD Training and Sample Size
- Measurements of Three-Level Hierarchical Structure in the Outliers in the Spectrum of Deepnet Hessians
- The Implicit and Explicit Regularization Effects of Dropout
- A Diffusion Theory For Deep Learning Dynamics: Stochastic Gradient Descent Exponentially Favors Flat Minima
- On the training dynamics of deep networks with regularization
- Deep learning versus kernel learning: an empirical study of loss landscape geometry and the time evolution of the Neural Tangent Kernel
- Emergent properties of the local geometry of neural loss landscapes
- Weight-space symmetry in deep networks gives rise to permutation saddles, connected by equal-loss valleys across the loss landscape
- Hessian-based toolbox for reliable and interpretable machine learning in physics
- Geometry of the Loss Landscape in Overparameterized Neural Networks: Symmetries and Invariances
- Wide-minima Density Hypothesis and the Explore-Exploit Learning Rate Schedule
- How noise affects the Hessian spectrum in overparameterized neural networks
- Sketching Curvature for Efficient Out-of-Distribution Detection for Deep Neural Networks
- Learning Rates as a Function of Batch Size: A Random Matrix Theory Approach to Neural Network Training
- SGD momentum optimizer with step estimation by online parabola model
- Pathological spectra of the Fisher information metric and its variants in deep neural networks
- Traces of Class/Cross-Class Structure Pervade Deep Learning Spectra
- Average-case Acceleration Through Spectral Density Estimation
- Interpreting Deep Learning: The Machine Learning Rorschach Test?
- Stability and Generalization of Stochastic Gradient Methods for Minimax Problems
- A Parsimonious Tour of Bayesian Model Uncertainty
- End-to-end Learning of a Convolutional Neural Network via Deep Tensor Decomposition
- Hessian based analysis of SGD for Deep Nets: Dynamics and Generalization
- Beyond Random Matrix Theory for Deep Networks
- Hessian Eigenspectra of More Realistic Nonlinear Models
- The asymptotic spectrum of the Hessian of DNN throughout training
- Convergence rates of Gibbs measures with degenerate minimum
- Analytic Insights into Structure and Rank of Neural Network Hessian Maps
- A Theory-Driven Self-Labeling Refinement Method for Contrastive Representation Learning
- Learning with Gradient Descent and Weakly Convex Losses
- The Limiting Dynamics of SGD: Modified Loss, Phase Space Oscillations, and Anomalous Diffusion
- Sparse Flows: Pruning Continuous-depth Models
- Implicit Gradient Regularization
- Analytic Characterization of the Hessian in Shallow ReLU Models: A Tale of Symmetry
- De-randomized PAC-Bayes Margin Bounds: Applications to Non-convex and Non-smooth Predictors
- Information-Theoretic Local Minima Characterization and Regularization
- Lifelong Learning with Sketched Structural Regularization
- On the Promise of the Stochastic Generalized Gauss-Newton Method for Training DNNs
- Approximate Newton-based statistical inference using only stochastic gradients
- Revisiting the Fragility of Influence Functions
- Cockpit: A Practical Debugging Tool for the Training of Deep Neural Networks
- Deep Curvature Suite
- Flatness is a False Friend
- How many degrees of freedom do we need to train deep networks: a loss landscape perspective
- What training reveals about neural network complexity
- On the Bias-Variance Tradeoff: Textbooks Need an Update
- Disentangling the Gauss-Newton Method and Approximate Inference for Neural Networks
- Structured Dropout Variational Inference for Bayesian Neural Networks
- Subaging in underparametrized Deep Neural Networks
- Appearance of Random Matrix Theory in Deep Learning
- Generalisation under gradient descent via deterministic PAC-Bayes
- Semiparametric Nonlinear Bipartite Graph Representation Learning with Provable Guarantees
- Sparsification as a Remedy for Staleness in Distributed Asynchronous SGD
- GENNI: Visualising the Geometry of Equivalences for Neural Network Identifiability
- Periodic Spectral Ergodicity: A Complexity Measure for Deep Neural Networks and Neural Architecture Search
- Analytic Study of Families of Spurious Minima in Two-Layer ReLU Neural Networks: A Tale of Symmetry II
- Toward Communication Efficient Adaptive Gradient Method
- Dynamic Game Theoretic Neural Optimizer
- Regularization in ResNet with Stochastic Depth