YellowFin and the Art of Momentum Tuning
arXiv:1706.03471
Abstract
Hyperparameter tuning is one of the most time-consuming workloads in deep learning. State-of-the-art optimizers, such as AdaGrad, RMSProp and Adam, reduce this labor by adaptively tuning an individual learning rate for each variable. Recently researchers have shown renewed interest in simpler methods like momentum SGD as they may yield better test metrics. Motivated by this trend, we ask: can simple adaptive methods based on SGD perform as well or better? We revisit the momentum SGD algorithm and show that hand-tuning a single learning rate and momentum makes it competitive with Adam. We then analyze its robustness to learning rate misspecification and objective curvature variation. Based on these insights, we design YellowFin, an automatic tuner for momentum and learning rate in SGD. YellowFin optionally uses a negative-feedback loop to compensate for the momentum dynamics in asynchronous settings on the fly. We empirically show that YellowFin can converge in fewer iterations than Adam on ResNets and LSTMs for image recognition, language modeling and constituency parsing, with a speedup of up to 3.28x in synchronous and up to 2.69x in asynchronous settings.
Updated to reflect improved stability discussion and work for SysML presentation
Cited by in corpus (19)
- Demystifying Parallel and Distributed Deep Learning: An In-Depth Concurrency Analysis
- Cooperative SGD: A unified Framework for the Design and Analysis of Communication-Efficient SGD Algorithms
- The Robust Manifold Defense: Adversarial Training using Generative Models
- Quasi-hyperbolic momentum and Adam for deep learning
- Descending through a Crowded Valley - Benchmarking Deep Learning Optimizers
- On the adequacy of untuned warmup for adaptive optimization
- Negative Momentum for Improved Game Dynamics
- Where Did My Optimum Go?: An Empirical Analysis of Gradient Descent Optimization in Policy Gradient Methods
- The Convergence of Sparsified Gradient Methods
- Taming Momentum in a Distributed Asynchronous Environment
- The Importance of Being Recurrent for Modeling Hierarchical Structure
- Reducing the variance in online optimization by transporting past gradients
- CROSSBOW: Scaling Deep Learning with Small Batch Sizes on Multi-GPU Servers
- SparCML: High-Performance Sparse Communication for Machine Learning
- The Convergence of Stochastic Gradient Descent in Asynchronous Shared Memory
- Asynchronous Optimization Methods for Efficient Training of Deep Neural Networks with Guarantees
- A Study of Condition Numbers for First-Order Optimization
- Backtracking gradient descent method for general functions, with applications to Deep Learning
- Robust Implicit Backpropagation