A Tail-Index Analysis of Stochastic Gradient Noise in Deep Neural Networks
arXiv:1901.06053
Abstract
The gradient noise (GN) in the stochastic gradient descent (SGD) algorithm is often considered to be Gaussian in the large data regime by assuming that the classical central limit theorem (CLT) kicks in. This assumption is often made for mathematical convenience, since it enables SGD to be analyzed as a stochastic differential equation (SDE) driven by a Brownian motion. We argue that the Gaussianity assumption might fail to hold in deep learning settings and hence render the Brownian motion-based analyses inappropriate. Inspired by non-Gaussian natural phenomena, we consider the GN in a more general context and invoke the generalized CLT (GCLT), which suggests that the GN converges to a heavy-tailed -stable random variable. Accordingly, we propose to analyze SGD as an SDE driven by a Lévy motion. Such SDEs can incur `jumps', which force the SDE transition from narrow minima to wider minima, as proven by existing metastability theory. To validate the -stable assumption, we conduct extensive experiments on common deep learning architectures and show that in all settings, the GN is highly non-Gaussian and admits heavy-tails. We further investigate the tail behavior in varying network architectures and sizes, loss functions, and datasets. Our results open up a different perspective and shed more light on the belief that SGD prefers wide minima.
Cited by in corpus (36)
- Review: Deep Learning in Electron Microscopy
- Towards Theoretically Understanding Why SGD Generalizes Better Than ADAM in Deep Learning
- The Heavy-Tail Phenomenon in SGD
- Stochasticity helps to navigate rough landscapes: comparing gradient-descent-based algorithms in the phase retrieval problem
- How Good is the Bayes Posterior in Deep Neural Networks Really?
- A Diffusion Theory For Deep Learning Dynamics: Stochastic Gradient Descent Exponentially Favors Flat Minima
- Non-Gaussianity of Stochastic Gradient Noise
- Fractional Underdamped Langevin Dynamics: Retargeting SGD with Momentum under Heavy-Tailed Gradient Noise
- On the Heavy-Tailed Theory of Stochastic Gradient Descent for Deep Neural Networks
- Coherent Gradients: An Approach to Understanding Generalization in Gradient Descent-based Optimization
- On the Validity of Modeling SGD with Stochastic Differential Equations (SDEs)
- A Study of Gradient Variance in Deep Learning
- Stochastic Optimization with Heavy-Tailed Noise via Accelerated Gradient Clipping
- First Exit Time Analysis of Stochastic Gradient Descent Under Heavy-Tailed Gradient Noise
- Positive-Negative Momentum: Manipulating Stochastic Gradient Noise to Improve Generalization
- Strength of Minibatch Noise in SGD
- Dynamic of Stochastic Gradient Descent with State-Dependent Noise
- Taxonomizing local versus global structure in neural network loss landscapes
- Stochastic Training is Not Necessary for Generalization
- The Limiting Dynamics of SGD: Modified Loss, Phase Space Oscillations, and Anomalous Diffusion
- Intrinsic Dimension, Persistent Homology and Generalization in Neural Networks
- Noise and Fluctuation of Finite Learning Rate Stochastic Gradient Descent
- High-probability Bounds for Non-Convex Stochastic Optimization with Heavy Tails
- An Empirical Study on the Intrinsic Privacy of SGD
- A Fully Spiking Hybrid Neural Network for Energy-Efficient Object Detection
- Approximating How Single Head Attention Learns
- On Proximal Policy Optimization's Heavy-tailed Gradients
- Asymmetric Heavy Tails and Implicit Bias in Gaussian Noise Injections
- Trap of Feature Diversity in the Learning of MLPs
- Fractional moment-preserving initialization schemes for training deep neural networks
- Revisiting the Characteristics of Stochastic Gradient Noise and Dynamics
- On the Sample Complexity and Metastability of Heavy-tailed Policy Search in Continuous Control
- Escaping Saddle Points with Stochastically Controlled Stochastic Gradient Methods
- Unique Properties of Flat Minima in Deep Networks
- A study on the plasticity of neural networks
- Training Efficiency and Robustness in Deep Learning