The asymptotic spectrum of the Hessian of DNN throughout training
arXiv:1910.02875
Abstract
The dynamics of DNNs during gradient descent is described by the so-called Neural Tangent Kernel (NTK). In this article, we show that the NTK allows one to gain precise insight into the Hessian of the cost of DNNs. When the NTK is fixed during training, we obtain a full characterization of the asymptotics of the spectrum of the Hessian, at initialization and during training. In the so-called mean-field limit, where the NTK is not fixed during training, we describe the first two moments of the Hessian at initialization.
References in corpus (7)
- The Loss Surfaces of Multilayer Networks
- Towards Understanding Generalization of Deep Learning: Perspective of Loss Landscapes
- Gradient Descent Happens in a Tiny Subspace
- An Investigation into Neural Net Optimization via Hessian Eigenvalue Density
- Mean-field theory of two-layers neural networks: dimension-free bounds and kernel limit
- On the saddle point problem for non-convex optimization
- Measurements of Three-Level Hierarchical Structure in the Outliers in the Spectrum of Deepnet Hessians