Zoneout: Regularizing RNNs by Randomly Preserving Hidden Activations
arXiv:1606.01305
Abstract
We propose zoneout, a novel method for regularizing RNNs. At each timestep, zoneout stochastically forces some hidden units to maintain their previous values. Like dropout, zoneout uses random noise to train a pseudo-ensemble, improving generalization. But by preserving instead of dropping hidden units, gradient information and state information are more readily propagated through time, as in feedforward stochastic depth networks. We perform an empirical investigation of various RNN regularizers, and find that zoneout gives significant performance improvements across tasks. We achieve competitive results with relatively simple models in character- and word-level language modelling on the Penn Treebank and Text8 datasets, and combining with recurrent batch normalization yields state-of-the-art results on permuted sequential MNIST.
David Krueger and Tegan Maharaj contributed equally to this work
References in corpus (6)
Cited by in corpus (80)
- An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling
- Gaussian Error Linear Units (GELUs)
- Recent Advances in Recurrent Neural Networks
- Style Tokens: Unsupervised Style Modeling, Control and Transfer in End-to-End Speech Synthesis
- Pointer Sentinel Mixture Models
- Regularizing and Optimizing LSTM Language Models
- Quasi-Recurrent Neural Networks
- Hierarchical Multiscale Recurrent Neural Networks
- One-Shot Imitation Learning
- Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions
- Flipout: Efficient Pseudo-Independent Weight Perturbations on Mini-Batches
- HiPPO: Recurrent Memory with Optimal Polynomial Projections
- An Analysis of Neural Language Modeling at Multiple Scales
- Independently Recurrent Neural Network (IndRNN): Building A Longer and Deeper RNN
- R-Transformer: Recurrent Neural Network Enhanced Transformer
- Multiplicative LSTM for sequence modelling
- ProphetNet: Predicting Future N-gram for Sequence-to-Sequence Pre-training
- Recurrent Independent Mechanisms
- The unreasonable effectiveness of the forget gate
- Augment your batch: better training with larger batches
- Learning End-to-end Multimodal Sensor Policies for Autonomous Navigation
- Survey of Dropout Methods for Deep Neural Networks
- Deep Convolutional Neural Network Design Patterns
- Memory Augmented Neural Networks with Wormhole Connections
- Revisiting Activation Regularization for Language RNNs
- Fast-Slow Recurrent Neural Networks
- Temporal Convolutional Attention-based Network For Sequence Modeling
- A global method to identify trees outside of closed-canopy forests with medium-resolution satellite imagery
- Skip RNN: Learning to Skip State Updates in Recurrent Neural Networks
- Dynamic Neural Turing Machine with Soft and Hard Addressing Schemes
- DropAttention: A Regularization Method for Fully-Connected Self-Attention Networks
- Improving the Gating Mechanism of Recurrent Neural Networks
- Learning to Create and Reuse Words in Open-Vocabulary Neural Language Modeling
- Regularizing RNNs for Caption Generation by Reconstructing The Past with The Present
- Neural Language Modeling by Jointly Learning Syntax and Lexicon
- Analyzing and Exploiting NARX Recurrent Neural Networks for Long-Term Dependencies
- Noisin: Unbiased Regularization for Recurrent Neural Networks
- Classification of Periodic Variable Stars with Novel Cyclic-Permutation Invariant Neural Networks
- Grow and Prune Compact, Fast, and Accurate LSTMs
- Bayesian Sparsification of Recurrent Neural Networks
- Discrete-Valued Neural Communication
- Improved Hierarchical Patient Classification with Language Model Pretraining over Clinical Notes
- Modularity in Deep Learning: A Survey
- Shifting Mean Activation Towards Zero with Bipolar Activation Functions
- Variational Bi-LSTMs
- Surprisal-Driven Zoneout
- Scheduled DropHead: A Regularization Method for Transformer Models
- Surprisal-Driven Feedback in Recurrent Networks
- Conversational End-to-End TTS for Voice Agent
- Effectiveness of Scaled Exponentially-Regularized Linear Units (SERLUs)
- Benchmarking Deep Sequential Models on Volatility Predictions for Financial Time Series
- Drop-Activation: Implicit Parameter Reduction and Harmonic Regularization
- WaveGrad 2: Iterative Refinement for Text-to-Speech Synthesis
- Semi-Supervised Training for Improving Data Efficiency in End-to-End Speech Synthesis
- Refined Gate: A Simple and Effective Gating Mechanism for Recurrent Units
- Video Representation Learning and Latent Concept Mining for Large-scale Multi-label Video Classification
- Adaptive Low-Rank Factorization to regularize shallow and deep neural networks
- Rotational Unit of Memory
- Shortcut Sequence Tagging
- Rapping-Singing Voice Synthesis based on Phoneme-level Prosody Control
- ResiliNet: Failure-Resilient Inference in Distributed Neural Networks
- Stochastic Weight Matrix-based Regularization Methods for Deep Neural Networks
- Cycle-consistency training for end-to-end speech recognition
- UniDrop: A Simple yet Effective Technique to Improve Transformer without Extra Cost
- Stochastic Bottleneck: Rateless Auto-Encoder for Flexible Dimensionality Reduction
- Learning Simpler Language Models with the Differential State Framework
- Recurrent Neural Network from Adder's Perspective: Carry-lookahead RNN
- Dropout with Tabu Strategy for Regularizing Deep Neural Networks
- Efficiently applying attention to sequential data with the Recurrent Discounted Attention unit
- Highway State Gating for Recurrent Highway Networks: improving information flow through time
- Regularizing Recurrent Neural Networks via Sequence Mixup
- Zero Training Overhead Portfolios for Learning to Solve Combinatorial Problems
- Reducing state updates via Gaussian-gated LSTMs
- ARMIN: Towards a More Efficient and Light-weight Recurrent Memory Network
- Gram Regularization for Multi-view 3D Shape Retrieval
- Adaptive Noise Injection: A Structure-Expanding Regularization for RNN
- Modeling of Rakugo Speech and Its Limitations: Toward Speech Synthesis That Entertains Audiences
- Meta-Forecasting by combining Global Deep Representations with Local Adaptation
- Adaptive Low-Rank Regularization with Damping Sequences to Restrict Lazy Weights in Deep Networks
- Structured in Space, Randomized in Time: Leveraging Dropout in RNNs for Efficient Training