Appearance of Random Matrix Theory in Deep Learning
arXiv:2102.06740 · doi:10.1016/j.physa.2021.126742
Abstract
We investigate the local spectral statistics of the loss surface Hessians of artificial neural networks, where we discover excellent agreement with Gaussian Orthogonal Ensemble statistics across several network architectures and datasets. These results shed new light on the applicability of Random Matrix Theory to modelling neural networks and suggest a previously unrecognised role for it in the study of loss surfaces in deep learning. Inspired by these observations, we propose a novel model for the true loss surfaces of neural networks, consistent with our observations, which allows for Hessian spectral densities with rank degeneracy and outliers, extensively observed in practice, and predicts a growing independence of loss gradients as a function of distance in weight-space. We further investigate the importance of the true loss surface in neural networks and find, in contrast to previous work, that the exponential hardness of locating the global minimum has practical consequences for achieving state of the art performance.
33 pages, 14 figures
References in corpus (12)
- The distribution of the ratio of consecutive level spacings in random matrix ensembles
- The Loss Surfaces of Multilayer Networks
- Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
- The Principles of Deep Learning Theory
- Cleaning large correlation matrices: tools from random matrix theory
- An Investigation into Neural Net Optimization via Hessian Eigenvalue Density
- Explorations on high dimensional landscapes
- The Neural Tangent Kernel in High Dimensions: Triple Descent and a Multi-Scale Theory of Generalization
- The Spectral Backbone of Excitation Transport in Ultra-Cold Rydberg Gases
- Measurements of Three-Level Hierarchical Structure in the Outliers in the Spectrum of Deepnet Hessians
- A Precise Performance Analysis of Learning with Random Features
- A spin-glass model for the loss surfaces of generative adversarial networks