A Theoretically Grounded Application of Dropout in Recurrent Neural Networks
arXiv:1512.05287
Abstract
Recurrent neural networks (RNNs) stand at the forefront of many recent developments in deep learning. Yet a major difficulty with these models is their tendency to overfit, with dropout shown to fail when applied to recurrent layers. Recent results at the intersection of Bayesian modelling and deep learning offer a Bayesian interpretation of common deep learning techniques such as dropout. This grounding of dropout in approximate Bayesian inference suggests an extension of the theoretical results, offering insights into the use of dropout with RNN models. We apply this new variational inference based dropout technique in LSTM and GRU models, assessing it on language modelling and sentiment analysis tasks. The new approach outperforms existing techniques, and to the best of our knowledge improves on the single model state-of-the-art in language modelling with the Penn Treebank (73.4 test perplexity). This extends our arsenal of variational tools in deep learning.
Added clarifications; Published in NIPS 2016
References in corpus (5)
Cited by in corpus (66)
- Neural Architecture Search with Reinforcement Learning
- An Introduction to Variational Autoencoders
- A trans-disciplinary review of deep learning research for water resources scientists
- DeepConv-DTI: Prediction of drug-target interactions via deep learning with convolution on protein sequences
- Convolutional Recurrent Neural Networks for Polyphonic Sound Event Detection
- Quantifying total uncertainty in physics-informed neural networks for solving forward and inverse stochastic problems
- Pointer Sentinel Mixture Models
- Light Gated Recurrent Units for Speech Recognition
- Extracting possibly representative COVID-19 Biomarkers from X-Ray images with Deep Learning approach and image data related to Pulmonary Diseases
- Deep and Confident Prediction for Time Series at Uber
- Accurate Uncertainties for Deep Learning Using Calibrated Regression
- Prolongation of SMAP to Spatio-temporally Seamless Coverage of Continental US Using a Deep Learning Neural Network
- An overview and comparative analysis of Recurrent Neural Networks for Short Term Load Forecasting
- Multimodal Residual Learning for Visual QA
- Recurrent Neural Networks for Multivariate Time Series with Missing Values
- Predicting Video Saliency with Object-to-Motion CNN and Two-layer Convolutional LSTM
- SuperNNova: an open-source framework for Bayesian, Neural Network based supernova classification
- Tying Word Vectors and Word Classifiers: A Loss Framework for Language Modeling
- Data Noising as Smoothing in Neural Network Language Models
- Causality Extraction based on Self-Attentive BiLSTM-CRF with Transferred Embeddings
- Learning the PE Header, Malware Detection with Minimal Domain Knowledge
- Learning a bidirectional mapping between human whole-body motion and natural language using deep recurrent neural networks
- On the Feasibility of Transfer-learning Code Smells using Deep Learning
- Deep Learning: A Bayesian Perspective
- A deep learning approach to cosmological dark energy models
- Recurrent Dropout without Memory Loss
- Advanced Dropout: A Model-free Methodology for Bayesian Dropout Optimization
- Robust and Subject-Independent Driving Manoeuvre Anticipation through Domain-Adversarial Recurrent Neural Networks
- Compression of Recurrent Neural Networks for Efficient Language Modeling
- End-to-End Dense Video Captioning with Masked Transformer
- An empirical study on the effectiveness of images in Multimodal Neural Machine Translation
- Learning Uncertainty with Artificial Neural Networks for Improved Predictive Process Monitoring
- Similarity Learning for Authorship Verification in Social Media
- Molecule Identification with Rotational Spectroscopy and Probabilistic Deep Learning
- PatchUp: A Feature-Space Block-Level Regularization Technique for Convolutional Neural Networks
- Uncertainty-Aware Attention for Reliable Interpretation and Prediction
- A Multi-Modal Chinese Poetry Generation Model
- Visualizing and Understanding Curriculum Learning for Long Short-Term Memory Networks
- Scientific Inference With Interpretable Machine Learning: Analyzing Models to Learn About Real-World Phenomena
- An Anomaly Detection Method for Satellites Using Monte Carlo Dropout
- Feature engineering vs. deep learning for paper section identification: Toward applications in Chinese medical literature
- Probabilistic Deep Learning to Quantify Uncertainty in Air Quality Forecasting
- Simple and Effective Knowledge-Driven Query Expansion for QA-Based Product Attribute Extraction
- Deep Factors with Gaussian Processes for Forecasting
- Bayesian LSTMs in medicine
- Quantifying Uncertainties in Natural Language Processing Tasks
- Actively Learning what makes a Discrete Sequence Valid
- DropDim: A Regularization Method for Transformer Networks
- Classification of Cardiac Arrhythmias from Single Lead ECG with a Convolutional Recurrent Neural Network
- Bayesian Sparsification of Recurrent Neural Networks
- Pushing the bounds of dropout
- Exploring Bayesian Deep Learning for Urgent Instructor Intervention Need in MOOC Forums
- A Way out of the Odyssey: Analyzing and Combining Recent Insights for LSTMs
- Reconstructing the Hubble diagram of gamma-ray bursts using deep learning
- Visually Grounded Word Embeddings and Richer Visual Features for Improving Multimodal Neural Machine Translation
- Cardiac Arrhythmia Detection from ECG with Convolutional Recurrent Neural Networks
- Convolution Aware Initialization
- Low-rank passthrough neural networks
- Delving Deeper into the Decoder for Video Captioning
- A Bayesian Machine Learning Algorithm for Predicting ENSO Using Short Observational Time Series
- Variational Smoothing in Recurrent Neural Network Language Models
- Eco-PiNN: A Physics-informed Neural Network for Eco-toll Estimation
- Naive imputation implicitly regularizes high-dimensional linear models
- Persistence pays off: Paying Attention to What the LSTM Gating Mechanism Persists
- Rapid Identification of X-ray Diffraction Spectra Based on Very Limited Data by Interpretable Convolutional Neural Networks
- Multi-talker Speech Separation with Utterance-level Permutation Invariant Training of Deep Recurrent Neural Networks