Disentangling Adaptive Gradient Methods from Learning Rates
arXiv:2002.11803
Abstract
We investigate several confounding factors in the evaluation of optimization algorithms for deep learning. Primarily, we take a deeper look at how adaptive gradient methods interact with the learning rate schedule, a notoriously difficult-to-tune hyperparameter which has dramatic effects on the convergence and generalization of neural network training. We introduce a "grafting" experiment which decouples an update's magnitude from its direction, finding that many existing beliefs in the literature may have arisen from insufficient isolation of the implicit schedule of step sizes. Alongside this contribution, we present some empirical and theoretical retrospectives on the generalization of adaptive gradient methods, aimed at bringing more clarity to this space.
References in corpus (12)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- ADADELTA: An Adaptive Learning Rate Method
- On the difficulty of training Recurrent Neural Networks
- On the Convergence of Adam and Beyond
- One weird trick for parallelizing convolutional neural networks
- Large Batch Training of Convolutional Networks
- Improving Generalization Performance by Switching from Adam to SGD
- DyNet: The Dynamic Neural Network Toolkit
- Fine-Grained Analysis of Optimization and Generalization for Overparameterized Two-Layer Neural Networks
- Lingvo: a Modular and Scalable Framework for Sequence-to-Sequence Modeling
- An Exponential Learning Rate Schedule for Deep Learning
- Span-Based Constituency Parsing with a Structure-Label System and Provably Optimal Dynamic Oracles