MomentumRNN: Integrating Momentum into Recurrent Neural Networks
arXiv:2006.06919
Abstract
Designing deep neural networks is an art that often involves an expensive search over candidate architectures. To overcome this for recurrent neural nets (RNNs), we establish a connection between the hidden state dynamics in an RNN and gradient descent (GD). We then integrate momentum into this framework and propose a new family of RNNs, called {\em MomentumRNNs}. We theoretically prove and numerically demonstrate that MomentumRNNs alleviate the vanishing gradient issue in training RNNs. We study the momentum long-short term memory (MomentumLSTM) and verify its advantages in convergence speed and accuracy over its LSTM counterpart across a variety of benchmarks. We also demonstrate that MomentumRNN is applicable to many types of recurrent cells, including those in the state-of-the-art orthogonal RNNs. Finally, we show that other advanced momentum-based optimization methods, such as Adam and Nesterov accelerated gradients with a restart, can be easily incorporated into the MomentumRNN framework for designing new recurrent cells with even better performance. The code is available at https://github.com/minhtannguyen/MomentumRNN.
21 pages, 11 figures, Accepted for publication at Advances in Neural Information Processing Systems (NeurIPS) 2020
References in corpus (9)
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- On the difficulty of training Recurrent Neural Networks
- Momentum Contrast for Unsupervised Visual Representation Learning
- A Simple Way to Initialize Recurrent Networks of Rectified Linear Units
- AntisymmetricRNN: A Dynamical System View on Recurrent Neural Networks
- Symplectic Recurrent Neural Networks
- Recurrent Neural Networks in the Eye of Differential Equations
- A recurrent neural network without chaos
- Understanding the Learned Iterative Soft Thresholding Algorithm with matrix factorization