Skip RNN: Learning to Skip State Updates in Recurrent Neural Networks
arXiv:1708.06834
Abstract
Recurrent Neural Networks (RNNs) continue to show outstanding performance in sequence modeling tasks. However, training RNNs on long sequences often face challenges like slow inference, vanishing gradients and difficulty in capturing long term dependencies. In backpropagation through time settings, these issues are tightly coupled with the large, sequential computational graph resulting from unfolding the RNN in time. We introduce the Skip RNN model which extends existing RNN models by learning to skip state updates and shortens the effective size of the computational graph. This model can also be encouraged to perform fewer state updates through a budget constraint. We evaluate the proposed model on various tasks and show how it can reduce the number of required RNN updates while preserving, and sometimes even improving, the performance of the baseline RNN models. Source code is publicly available at https://imatge-upc.github.io/skiprnn-2017-telecombcn/ .
Accepted as conference paper at ICLR 2018
References in corpus (14)
- Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling
- Two-Stream Convolutional Networks for Action Recognition in Videos
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation
- Recurrent Neural Network Regularization
- Recurrent Models of Visual Attention
- Multiple Object Recognition with Visual Attention
- A Simple Way to Initialize Recurrent Networks of Rectified Linear Units
- Quasi-Recurrent Neural Networks
- Hierarchical Multiscale Recurrent Neural Networks
- Phased LSTM: Accelerating Recurrent Network Training for Long or Event-based Sequences
- Variable Computation in Recurrent Neural Networks
- A Way out of the Odyssey: Analyzing and Combining Recent Insights for LSTMs
- Asynchronous Temporal Fields for Action Recognition
Cited by in corpus (16)
- EleAtt-RNN: Adding Attentiveness to Neurons in Recurrent Neural Networks
- Neural Speed Reading via Skim-RNN
- Neural Rough Differential Equations for Long Time Series
- Focused Hierarchical RNNs for Conditional Sequence Processing
- Cross-Modality Attention with Semantic Graph Embedding for Multi-Label Classification
- Switchable Precision Neural Networks
- Deep Learning for Multi-Scale Changepoint Detection in Multivariate Time Series
- Benchmarking Deep Sequential Models on Volatility Predictions for Financial Time Series
- Equilibrated Recurrent Neural Network: Neuronal Time-Delayed Self-Feedback Improves Accuracy and Stability
- Recognizing Video Events with Varying Rhythms
- Long Short-Term Memory with Dynamic Skip Connections
- Learning to Adaptively Scale Recurrent Neural Networks
- Sibling Neural Estimators: Improving Iterative Image Decoding with Gradient Communication
- RNN Training along Locally Optimal Trajectories via Frank-Wolfe Algorithm
- Adversarially Robust and Explainable Model Compression with On-Device Personalization for Text Classification
- ARMIN: Towards a More Efficient and Light-weight Recurrent Memory Network