The Full Spectrum of Deepnet Hessians at Scale: Dynamics with SGD Training and Sample Size
arXiv:1811.07062
Abstract
We apply state-of-the-art tools in modern high-dimensional numerical linear algebra to approximate efficiently the spectrum of the Hessian of modern deepnets, with tens of millions of parameters, trained on real data. Our results corroborate previous findings, based on small-scale networks, that the Hessian exhibits "spiked" behavior, with several outliers isolated from a continuous bulk. We decompose the Hessian into different components and study the dynamics with training and sample size of each term individually.
References in corpus (12)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Fashion-MNIST: a Novel Image Dataset for Benchmarking Machine Learning Algorithms
- Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
- Three Factors Influencing Minima in SGD
- An Empirical Model of Large-Batch Training
- Gradient Descent Happens in a Tiny Subspace
- Implicit Self-Regularization in Deep Neural Networks: Evidence from Random Matrix Theory and Implications for Learning
- An Investigation into Neural Net Optimization via Hessian Eigenvalue Density
- Hessian-based Analysis of Large Batch Training and Robustness to Adversaries
- Fluctuation-dissipation relations for stochastic gradient descent
- Measurements of Three-Level Hierarchical Structure in the Outliers in the Spectrum of Deepnet Hessians
- A Tail-Index Analysis of Stochastic Gradient Noise in Deep Neural Networks
Cited by in corpus (25)
- Rethinking Parameter Counting in Deep Models: Effective Dimensionality Revisited
- Mathematical Models of Overparameterized Neural Networks
- Recent advances in deep learning theory
- On the Heavy-Tailed Theory of Stochastic Gradient Descent for Deep Neural Networks
- The Loss Surfaces of Neural Networks with General Activation Functions
- Directional Pruning of Deep Neural Networks
- Dissecting Hessian: Understanding Common Structure of Hessian in Neural Networks
- A spin-glass model for the loss surfaces of generative adversarial networks
- Learning Rates as a Function of Batch Size: A Random Matrix Theory Approach to Neural Network Training
- The Geometry of Sign Gradient Descent
- Analysis of stochastic Lanczos quadrature for spectrum approximation
- Beyond Random Matrix Theory for Deep Networks
- A Loss Curvature Perspective on Training Instability in Deep Learning
- Exploring Weight Importance and Hessian Bias in Model Pruning
- The Limiting Dynamics of SGD: Modified Loss, Phase Space Oscillations, and Anomalous Diffusion
- SGD in the Large: Average-case Analysis, Asymptotics, and Stepsize Criticality
- Analytic Characterization of the Hessian in Shallow ReLU Models: A Tale of Symmetry
- De-randomized PAC-Bayes Margin Bounds: Applications to Non-convex and Non-smooth Predictors
- Deep Curvature Suite
- Does the Data Induce Capacity Control in Deep Learning?
- Flatness is a False Friend
- Appearance of Random Matrix Theory in Deep Learning
- ViViT: Curvature access through the generalized Gauss-Newton's low-rank structure
- Analytic Study of Families of Spurious Minima in Two-Layer ReLU Neural Networks: A Tale of Symmetry II
- Dynamics of Stochastic Momentum Methods on Large-scale, Quadratic Models