Fixup Initialization: Residual Learning Without Normalization
arXiv:1901.09321
Abstract
Normalization layers are a staple in state-of-the-art deep neural network architectures. They are widely believed to stabilize training, enable higher learning rate, accelerate convergence and improve generalization, though the reason for their effectiveness is still an active research topic. In this work, we challenge the commonly-held beliefs by showing that none of the perceived benefits is unique to normalization. Specifically, we propose fixed-update initialization (Fixup), an initialization motivated by solving the exploding and vanishing gradient problem at the beginning of training via properly rescaling a standard initialization. We find training residual networks with Fixup to be as stable as training with normalization -- even for networks with 10,000 layers. Furthermore, with proper regularization, Fixup enables residual networks without normalization to achieve state-of-the-art performance in image classification and machine translation.
Updating reference. Accepted for publication at ICLR 2019; see https://openreview.net/forum?id=H1gsz30cKX
References in corpus (4)
Cited by in corpus (29)
- Massively Multilingual Neural Machine Translation in the Wild: Findings and Challenges
- Reducing Transformer Depth on Demand with Structured Dropout
- High-Performance Large-Scale Image Recognition Without Normalization
- Optimization for deep learning: theory and algorithms
- Understanding and Improving Transformer From a Multi-Particle Dynamic System Point of View
- Jukebox: A Generative Model for Music
- Measuring the Algorithmic Efficiency of Neural Networks
- Explicit Sparse Transformer: Concentrated Attention Through Explicit Selection
- Very Deep VAEs Generalize Autoregressive Models and Can Outperform Them on Images
- MUSE: Parallel Multi-Scale Attention for Sequence to Sequence Learning
- Towards Efficient Training for Neural Network Quantization
- An Adaptive and Momental Bound Method for Stochastic Learning
- Orthogonalizing Convolutional Layers with the Cayley Transform
- Scaling Imitation Learning in Minecraft
- Analyzing Monotonic Linear Interpolation in Neural Network Loss Landscapes
- A Loss Curvature Perspective on Training Instability in Deep Learning
- Deep Connectomics Networks: Neural Network Architectures Inspired by Neuronal Networks
- Mean Shift Rejection: Training Deep Neural Networks Without Minibatch Statistics or Normalization
- Parameter Prediction for Unseen Deep Architectures
- How many winning tickets are there in one DNN?
- On the Demystification of Knowledge Distillation: A Residual Network Perspective
- Comparing the costs of abstraction for DL frameworks
- Free-viewpoint Indoor Neural Relighting from Multi-view Stereo
- "BNN - BN = ?": Training Binary Neural Networks without Batch Normalization
- Self-Attentive Ensemble Transformer: Representing Ensemble Interactions in Neural Networks for Earth System Models
- Farkas layers: don't shift the data, fix the geometry
- On the generalization of bayesian deep nets for multi-class classification
- Seven Myths in Machine Learning Research
- Modeling from Features: a Mean-field Framework for Over-parameterized Deep Neural Networks