Recurrent Dropout without Memory Loss
arXiv:1603.05118
Abstract
This paper presents a novel approach to recurrent neural network (RNN) regularization. Differently from the widely adopted dropout method, which is applied to \textit{forward} connections of feed-forward architectures or RNNs, we propose to drop neurons directly in \textit{recurrent} connections in a way that does not cause loss of long-term memory. Our approach is as easy to implement and apply as the regular feed-forward dropout and we demonstrate its effectiveness for Long Short-Term Memory network, the most popular type of RNN cells. Our experiments on NLP benchmarks show consistent improvements even when combined with conventional feed-forward dropout.
References in corpus (7)
- Sequence to Sequence Learning with Neural Networks
- Improving neural networks by preventing co-adaptation of feature detectors
- Recurrent Neural Network Regularization
- On the Properties of Neural Machine Translation: Encoder-Decoder Approaches
- A Theoretically Grounded Application of Dropout in Recurrent Neural Networks
- Transition-Based Dependency Parsing with Stack Long Short-Term Memory
- Learning Longer Memory in Recurrent Neural Networks
Cited by in corpus (43)
- Recent Advances in Recurrent Neural Networks
- Prolongation of SMAP to Spatio-temporally Seamless Coverage of Continental US Using a Deep Learning Neural Network
- Phased LSTM: Accelerating Recurrent Network Training for Long or Event-based Sequences
- Data Noising as Smoothing in Neural Network Language Models
- Zoneout: Regularizing RNNs by Randomly Preserving Hidden Activations
- Independently Recurrent Neural Network (IndRNN): Building A Longer and Deeper RNN
- The Sockeye 2 Neural Machine Translation Toolkit at AMTA 2020
- Dual Rectified Linear Units (DReLUs): A Replacement for Tanh Activation Functions in Quasi-Recurrent Neural Networks
- Dynamic Neural Turing Machine with Soft and Hard Addressing Schemes
- DropAttention: A Regularization Method for Fully-Connected Self-Attention Networks
- Few-Shot Generalization Across Dialogue Tasks
- Predicting Future Lane Changes of Other Highway Vehicles using RNN-based Deep Models
- AMPNet: Asynchronous Model-Parallel Training for Dynamic Neural Networks
- Multilingual Training and Cross-lingual Adaptation on CTC-based Acoustic Model
- The Implicit and Explicit Regularization Effects of Dropout
- Learning to Create and Reuse Words in Open-Vocabulary Neural Language Modeling
- Improving the Neural GPU Architecture for Algorithm Learning
- Hierarchical Temporal Convolutional Networks for Dynamic Recommender Systems
- Noisin: Unbiased Regularization for Recurrent Neural Networks
- Learning to recognize touch gestures: recurrent vs. convolutional features and dynamic sampling
- Bayesian Sparsification of Recurrent Neural Networks
- Scheduling Computation Graphs of Deep Learning Models on Manycore CPUs
- Pushing the bounds of dropout
- Predicting Movie Genres Based on Plot Summaries
- Shifting Mean Activation Towards Zero with Bipolar Activation Functions
- Improving LSTM-CTC based ASR performance in domains with limited training data
- Almost Sure Convergence of Dropout Algorithms for Neural Networks
- Scalable Bayesian Learning of Recurrent Neural Networks for Language Modeling
- Divide and Conquer: A Deep CASA Approach to Talker-independent Monaural Speaker Separation
- Recurrent Memory Array Structures
- Dependency Parsing as Head Selection
- Low-Rank RNN Adaptation for Context-Aware Language Modeling
- DTMT: A Novel Deep Transition Architecture for Neural Machine Translation
- Learning to Adaptively Scale Recurrent Neural Networks
- Fine-tuning Handwriting Recognition systems with Temporal Dropout
- Long Short-Term Network Based Unobtrusive Perceived Workload Monitoring with Consumer Grade Smartwatches in the Wild
- Adversarial Dropout for Recurrent Neural Networks
- Learning to Compose over Tree Structures via POS Tags
- Subword Language Model for Query Auto-Completion
- Medi-Care AI: Predicting Medications From Billing Codes via Robust Recurrent Neural Networks
- Structured in Space, Randomized in Time: Leveraging Dropout in RNNs for Efficient Training
- How Do Neural Sequence Models Generalize? Local and Global Context Cues for Out-of-Distribution Prediction
- Thick-Net: Parallel Network Structure for Sequential Modeling