Understanding the Role of Momentum in Stochastic Gradient Methods
arXiv:1910.13962
Abstract
The use of momentum in stochastic gradient methods has become a widespread practice in machine learning. Different variants of momentum, including heavy-ball momentum, Nesterov's accelerated gradient (NAG), and quasi-hyperbolic momentum (QHM), have demonstrated success on various tasks. Despite these empirical successes, there is a lack of clear understanding of how the momentum parameters affect convergence and various performance measures of different algorithms. In this paper, we use the general formulation of QHM to give a unified analysis of several popular algorithms, covering their asymptotic convergence conditions, stability regions, and properties of their stationary distributions. In addition, by combining the results on convergence rates and stationary distributions, we obtain sometimes counter-intuitive practical guidelines for setting the learning rate and momentum parameters.
33rd Conference on Neural Information Processing Systems (NeurIPS 2019), Vancouver, Canada
Cited by in corpus (18)
- SlowMo: Improving Communication-Efficient Distributed SGD with Slow Momentum
- Quasi-Global Momentum: Accelerating Decentralized Deep Learning on Heterogeneous Data
- An Improved Analysis of Stochastic Gradient Descent with Momentum
- A Selective Review on Statistical Methods for Massive Data Computation: Distributed Computing, Subsampling, and Minibatch Techniques
- Communication Efficient Distributed Learning with Censored, Quantized, and Generalized Group ADMM
- Strength of Minibatch Noise in SGD
- The Role of Momentum Parameters in the Optimal Convergence of Adaptive Polyak's Heavy-ball Methods
- Gradient descent with momentum --- to accelerate or to super-accelerate?
- Training Deep Neural Networks with Adaptive Momentum Inspired by the Quadratic Optimization
- Noise and Fluctuation of Finite Learning Rate Stochastic Gradient Descent
- An Asymptotic Analysis of Minibatch-Based Momentum Methods for Linear Regression Models
- Analytical Study of Momentum-Based Acceleration Methods in Paradigmatic High-Dimensional Non-Convex Problems
- Quickly Finding a Benign Region via Heavy Ball Momentum in Non-Convex Optimization
- Bandwidth-based Step-Sizes for Non-Convex Stochastic Optimization
- Revisiting the Role of Euler Numerical Integration on Acceleration and Stability in Convex Optimization
- Understanding Modern Techniques in Optimization: Frank-Wolfe, Nesterov's Momentum, and Polyak's Momentum
- Scaling transition from momentum stochastic gradient descent to plain stochastic gradient descent
- A Modular Analysis of Provable Acceleration via Polyak's Momentum: Training a Wide ReLU Network and a Deep Linear Network