Stable Recurrent Models
arXiv:1805.10369
Abstract
Stability is a fundamental property of dynamical systems, yet to this date it has had little bearing on the practice of recurrent neural networks. In this work, we conduct a thorough investigation of stable recurrent models. Theoretically, we prove stable recurrent neural networks are well approximated by feed-forward networks for the purpose of both inference and training by gradient descent. Empirically, we demonstrate stable recurrent models often perform as well as their unstable counterparts on benchmark sequence tasks. Taken together, these findings shed light on the effective power of recurrent networks and suggest much of sequence learning happens, or can be made to happen, in the stable regime. Moreover, our results help to explain why in many cases practitioners succeed in replacing recurrent models by feed-forward models.
To appear in ICLR 2019. This paper was previously titled "When Recurrent Models Don't Need to Be Recurrent." The current version subsumes all previous versions
References in corpus (6)
- An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling
- On the difficulty of training Recurrent Neural Networks
- WaveNet: A Generative Model for Raw Audio
- Full-Capacity Unitary Recurrent Neural Networks
- Non-Asymptotic Analysis of Robust Control from Coarse-Grained Identification
- A recurrent neural network without chaos
Cited by in corpus (32)
- Theoretical Limitations of Self-Attention in Neural Sequence Models
- Trellis Networks for Sequence Modeling
- On the stability properties of Gated Recurrent Units neural networks
- Beyond exploding and vanishing gradients: analysing RNN training using attractors and smoothness
- Lipschitz Recurrent Neural Networks
- DeepFD: Automated Fault Diagnosis and Localization for Deep Learning Programs
- Stability of discrete-time feed-forward neural networks in NARX configuration
- Forecasting Sequential Data using Consistent Koopman Autoencoders
- Recurrent Neural Network-based Internal Model Control design for stable nonlinear systems
- LSTM Neural Networks: Input to State Stability and Probabilistic Safety Verification
- Non-asymptotic and Accurate Learning of Nonlinear Dynamical Systems
- Learning to Remember More with Less Memorization
- Learning 3D Human Dynamics from Video
- Almost Surely Stable Deep Dynamics
- Stable and expressive recurrent vision models
- Noisy Recurrent Neural Networks
- Convex Programming for Estimation in Nonlinear Recurrent Models
- Universal Approximation of Input-Output Maps by Temporal Convolutional Nets
- Deep learning for pedestrians: backpropagation in CNNs
- Equilibrated Recurrent Neural Network: Neuronal Time-Delayed Self-Feedback Improves Accuracy and Stability
- RNNs Evolving on an Equilibrium Manifold: A Panacea for Vanishing and Exploding Gradients?
- Active Deep Learning on Entity Resolution by Risk Sampling
- On the Privacy Risks of Deploying Recurrent Neural Networks in Machine Learning Models
- Learning Dynamics Models with Stable Invariant Sets
- RNN-based Online Learning: An Efficient First-Order Optimization Algorithm with a Convergence Guarantee
- Near-optimal Offline and Streaming Algorithms for Learning Non-Linear Dynamical Systems
- Convergent Graph Solvers
- Memory and attention in deep learning
- Using holistic event information in the trigger
- Modeling Electrical Motor Dynamics using Encoder-Decoder with Recurrent Skip Connection
- RNN Training along Locally Optimal Trajectories via Frank-Wolfe Algorithm
- Rényi Divergence in General Hidden Markov Models