Recurrent Batch Normalization
arXiv:1603.09025
Abstract
We propose a reparameterization of LSTM that brings the benefits of batch normalization to recurrent neural networks. Whereas previous works only apply batch normalization to the input-to-hidden transformation of RNNs, we demonstrate that it is both possible and beneficial to batch-normalize the hidden-to-hidden transition, thereby reducing internal covariate shift between time steps. We evaluate our proposal on various sequential problems such as sequence classification, language modeling and question answering. Our empirical results show that our batch-normalized LSTM consistently leads to faster convergence and improved generalization.
References in corpus (8)
- On the difficulty of training Recurrent Neural Networks
- Deep Speech 2: End-to-End Speech Recognition in English and Mandarin
- A Simple Way to Initialize Recurrent Networks of Rectified Linear Units
- Hierarchical Multiscale Recurrent Neural Networks
- Unitary Evolution Recurrent Neural Networks
- Bridging the Gaps Between Residual Learning, Recurrent Neural Networks and Visual Cortex
- Zoneout: Regularizing RNNs by Randomly Preserving Hidden Activations
- Architectural Complexity Measures of Recurrent Neural Networks
Cited by in corpus (38)
- Enhancing the Locality and Breaking the Memory Bottleneck of Transformer on Time Series Forecasting
- Zoneout: Regularizing RNNs by Randomly Preserving Hidden Activations
- Independently Recurrent Neural Network (IndRNN): Building A Longer and Deeper RNN
- Review: Deep Learning in Electron Microscopy
- Root Mean Square Layer Normalization
- Residual Connections Encourage Iterative Inference
- Decoupled Neural Interfaces using Synthetic Gradients
- Understanding Batch Normalization
- Sparse Attentive Backtracking: Temporal CreditAssignment Through Reminding
- Keys to Accurate Feature Extraction Using Residual Spiking Neural Networks
- Pediatric Automatic Sleep Staging: A comparative study of state-of-the-art deep learning methods
- Adaptive Regularization of Labels
- Hierarchical Temporal Convolutional Networks for Dynamic Recommender Systems
- Neural Language Modeling by Jointly Learning Syntax and Lexicon
- Batch Kalman Normalization: Towards Training Deep Neural Networks with Micro-Batches
- Analyzing and Exploiting NARX Recurrent Neural Networks for Long-Term Dependencies
- Noisin: Unbiased Regularization for Recurrent Neural Networks
- Bayesian LSTMs in medicine
- Techniques for visualizing LSTMs applied to electrocardiograms
- Low-Precision Batch-Normalized Activations
- Extended Batch Normalization
- ACtuAL: Actor-Critic Under Adversarial Learning
- LARNN: Linear Attention Recurrent Neural Network
- A dataset and exploration of models for understanding video data through fill-in-the-blank question-answering
- Attentive batch normalization for lstm-based acoustic modeling of speech recognition
- Towards Understanding Normalization in Neural ODEs
- Batch-normalized Recurrent Highway Networks
- SeqMobile: A Sequence Based Efficient Android Malware Detection System Using RNN on Mobile Devices
- Video Representation Learning and Latent Concept Mining for Large-scale Multi-label Video Classification
- Revisit Batch Normalization: New Understanding from an Optimization View and a Refinement via Composition Optimization
- Deep Triphone Embedding Improves Phoneme Recognition
- Accelerating Training of Deep Neural Networks with a Standardization Loss
- Learning to Adaptively Scale Recurrent Neural Networks
- Learning Simpler Language Models with the Differential State Framework
- Improving Gated Recurrent Unit Based Acoustic Modeling with Batch Normalization and Enlarged Context
- Double Forward Propagation for Memorized Batch Normalization
- Batch-normalized joint training for DNN-based distant speech recognition
- Adaptive Noise Injection: A Structure-Expanding Regularization for RNN