Layer Normalization
arXiv:1607.06450
Abstract
Training state-of-the-art, deep neural networks is computationally expensive. One way to reduce the training time is to normalize the activities of the neurons. A recently introduced technique called batch normalization uses the distribution of the summed input to a neuron over a mini-batch of training cases to compute a mean and variance which are then used to normalize the summed input to that neuron on each training case. This significantly reduces the training time in feed-forward neural networks. However, the effect of batch normalization is dependent on the mini-batch size and it is not obvious how to apply it to recurrent neural networks. In this paper, we transpose batch normalization into layer normalization by computing the mean and variance used for normalization from all of the summed inputs to the neurons in a layer on a single training case. Like batch normalization, we also give each neuron its own adaptive bias and gain which are applied after the normalization but before the non-linearity. Unlike batch normalization, layer normalization performs exactly the same computation at training and test times. It is also straightforward to apply to recurrent neural networks by computing the normalization statistics separately at each time step. Layer normalization is very effective at stabilizing the hidden state dynamics in recurrent networks. Empirically, we show that layer normalization can substantially reduce the training time compared with previously published techniques.
References in corpus (4)
Cited by in corpus (103)
- Searching for Activation Functions
- Artificial neural networks for neuroscientists: A primer
- SMASH: One-Shot Model Architecture Search through HyperNetworks
- Parameter Space Noise for Exploration
- Deep Closest Point: Learning Representations for Point Cloud Registration
- Local Light Field Fusion: Practical View Synthesis with Prescriptive Sampling Guidelines
- Beyond Self-attention: External Attention using Two Linear Layers for Visual Tasks
- Photographic Image Synthesis with Cascaded Refinement Networks
- Distance-based Self-Attention Network for Natural Language Inference
- FIGR: Few-shot Image Generation with Reptile
- Multimodal Entity Linking for Tweets
- A Hybrid Convolutional Variational Autoencoder for Text Generation
- Accelerating Training of Transformer-Based Language Models with Progressive Layer Dropping
- Revisiting Low-Resource Neural Machine Translation: A Case Study
- Online Learning for Neural Machine Translation Post-editing
- Pre-trained Language Model Representations for Language Generation
- Optimal Subarchitecture Extraction For BERT
- Multi-head or Single-head? An Empirical Comparison for Transformer Training
- From Here to There: Video Inbetweening Using Direct 3D Convolutions
- Disentangled Makeup Transfer with Generative Adversarial Network
- Convolutional Self-Attention Networks
- Associating Objects with Transformers for Video Object Segmentation
- Channel Normalization in Convolutional Neural Network avoids Vanishing Gradients
- Have convolutions already made recurrence obsolete for unconstrained handwritten text recognition ?
- Attend2Pack: Bin Packing through Deep Reinforcement Learning with Attention
- Adjusting for Dropout Variance in Batch Normalization and Weight Initialization
- Pre-Trained Models: Past, Present and Future
- The YouTube-8M Kaggle Competition: Challenges and Methods
- Projection Based Weight Normalization for Deep Neural Networks
- Keyphrase Extraction with Span-based Feature Representations
- Rethinking Skip Connection with Layer Normalization in Transformers and ResNets
- Unified Multi-Criteria Chinese Word Segmentation with BERT
- Towards Robust ResNet: A Small Step but A Giant Leap
- Preventing Posterior Collapse with delta-VAEs
- Capsule networks with non-iterative cluster routing
- Iterative Normalization: Beyond Standardization towards Efficient Whitening
- Multi-style Generative Reading Comprehension
- Spatial-Temporal Self-Attention Network for Flow Prediction
- Scalable variational Monte Carlo with graph neural ansatz
- Personalized Re-ranking for Recommendation
- Understanding and Improving Encoder Layer Fusion in Sequence-to-Sequence Learning
- Batch Group Normalization
- Language Model as an Annotator: Exploring DialoGPT for Dialogue Summarization
- Not All Attention Is All You Need
- Signal Combination for Language Identification
- Distilling Self-Knowledge From Contrastive Links to Classify Graph Nodes Without Passing Messages
- End-to-end Lane Shape Prediction with Transformers
- ISTD-GCN: Iterative Spatial-Temporal Diffusion Graph Convolutional Network for Traffic Speed Forecasting
- Intermediate Loss Regularization for CTC-based Speech Recognition
- Context Aware Machine Learning
- A Comparison of Approaches to Document-level Machine Translation
- Guided Transformer: Leveraging Multiple External Sources for Representation Learning in Conversational Search
- Exploring BERT Parameter Efficiency on the Stanford Question Answering Dataset v2.0
- Global-to-Local Neural Networks for Document-Level Relation Extraction
- Enhancing SAT solvers with glue variable predictions
- Making Neural Machine Reading Comprehension Faster
- Proportionate gradient updates with PercentDelta
- Hybrid Generative-Contrastive Representation Learning
- Modulating Image Restoration with Continual Levels via Adaptive Feature Modification Layers
- The Impact of Reinitialization on Generalization in Convolutional Neural Networks
- Neural Machine Translation with Joint Representation
- Find the Conversation Killers: a Predictive Study of Thread-ending Posts
- MOI-Mixer: Improving MLP-Mixer with Multi Order Interactions in Sequential Recommendation
- Regularizing Dialogue Generation by Imitating Implicit Scenarios
- Multi-Scale Self-Attention for Text Classification
- GlyphCRM: Bidirectional Encoder Representation for Chinese Character with its Glyph
- Towards Non-saturating Recurrent Units for Modelling Long-term Dependencies
- InSRL: A Multi-view Learning Framework Fusing Multiple Information Sources for Distantly-supervised Relation Extraction
- Discovering Useful Sentence Representations from Large Pretrained Language Models
- IART: Intent-aware Response Ranking with Transformers in Information-seeking Conversation Systems
- CasEE: A Joint Learning Framework with Cascade Decoding for Overlapping Event Extraction
- On the Role of Optimization in Double Descent: A Least Squares Study
- Semi-Autoregressive Transformer for Image Captioning
- On the Sub-Layer Functionalities of Transformer Decoder
- Approximated Orthonormal Normalisation in Training Neural Networks
- Diffusion models for Handwriting Generation
- A Transformer-based Math Language Model for Handwritten Math Expression Recognition
- RAMS-Trans: Recurrent Attention Multi-scale Transformer forFine-grained Image Recognition
- HySPA: Hybrid Span Generation for Scalable Text-to-Graph Extraction
- Exact-K Recommendation via Maximal Clique Optimization
- Incorporating Sememes into Chinese Definition Modeling
- Differentiable Sampling with Flexible Reference Word Order for Neural Machine Translation
- Self-attention Comparison Module for Boosting Performance on Retrieval-based Open-Domain Dialog Systems
- Linearly Constrained Weights: Reducing Activation Shift for Faster Training of Neural Networks
- Actions Generation from Captions
- Irregular Convolutional Auto-Encoder on Point Clouds
- Learning Propagation for Arbitrarily-structured Data
- Diverse Exploration via Conjugate Policies for Policy Gradient Methods
- Power Law Graph Transformer for Machine Translation and Representation Learning
- Network Horizon Dynamics I: Qualitative Aspects
- Marginal Utility Diminishes: Exploring the Minimum Knowledge for BERT Knowledge Distillation
- Stochastic Whitening Batch Normalization
- Generating Informative Dialogue Responses with Keywords-Guided Networks
- Keeping Notes: Conditional Natural Language Generation with a Scratchpad Mechanism
- New Interpretations of Normalization Methods in Deep Learning
- Riiid! Answer Correctness Prediction Kaggle Challenge: 4th Place Solution Summary
- 3D-MOV: Audio-Visual LSTM Autoencoder for 3D Reconstruction of Multiple Objects from Video
- Efficient Modelling Across Time of Human Actions and Interactions
- Controlling Neural Machine Translation Formality with Synthetic Supervision
- SQALER: Scaling Question Answering by Decoupling Multi-Hop and Logical Reasoning
- Tag and Correct: Question aware Open Information Extraction with Two-stage Decoding
- Spot What Matters: Learning Context Using Graph Convolutional Networks for Weakly-Supervised Action Detection
- Distribution Conditional Denoising: A Flexible Discriminative Image Denoiser