On the Distributional Properties of Adaptive Gradients
arXiv:2105.07222
Abstract
Adaptive gradient methods have achieved remarkable success in training deep neural networks on a wide variety of tasks. However, not much is known about the mathematical and statistical properties of this family of methods. This work aims at providing a series of theoretical analyses of its statistical properties justified by experiments. In particular, we show that when the underlying gradient obeys a normal distribution, the variance of the magnitude of the \textit{update} is an increasing and bounded function of time and does not diverge. This work suggests that the divergence of variance is not the cause of the need for warm up of the Adam optimizer, contrary to what is believed in the current literature.
References in corpus (11)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- On the Convergence of Adam and Beyond
- ALBERT: A Lite BERT for Self-supervised Learning of Language Representations
- On the Variance of the Adaptive Learning Rate and Beyond
- Adaptive Gradient Methods with Dynamic Bound of Learning Rate
- On Empirical Comparisons of Optimizers for Deep Learning
- Optimization for deep learning: theory and algorithms
- Theory of Deep Learning IIb: Optimization Properties of SGD
- The Heavy-Tail Phenomenon in SGD
- On the adequacy of untuned warmup for adaptive optimization
- LaProp: Separating Momentum and Adaptivity in Adam