Fixup Initialization: Residual Learning Without Normalization
arXiv:1901.09321
Abstract
Normalization layers are a staple in state-of-the-art deep neural network architectures. They are widely believed to stabilize training, enable higher learning rate, accelerate convergence and improve generalization, though the reason for their effectiveness is still an active research topic. In this work, we challenge the commonly-held beliefs by showing that none of the perceived benefits is unique to normalization. Specifically, we propose fixed-update initialization (Fixup), an initialization motivated by solving the exploding and vanishing gradient problem at the beginning of training via properly rescaling a standard initialization. We find training residual networks with Fixup to be as stable as training with normalization -- even for networks with 10,000 layers. Furthermore, with proper regularization, Fixup enables residual networks without normalization to achieve state-of-the-art performance in image classification and machine translation.
Updating reference. Accepted for publication at ICLR 2019; see https://openreview.net/forum?id=H1gsz30cKX
References in corpus (4)
Cited by in corpus (19)
- Massively Multilingual Neural Machine Translation in the Wild: Findings and Challenges
- Reducing Transformer Depth on Demand with Structured Dropout
- Optimization for deep learning: theory and algorithms
- Understanding and Improving Transformer From a Multi-Particle Dynamic System Point of View
- Jukebox: A Generative Model for Music
- Measuring the Algorithmic Efficiency of Neural Networks
- Explicit Sparse Transformer: Concentrated Attention Through Explicit Selection
- MUSE: Parallel Multi-Scale Attention for Sequence to Sequence Learning
- Towards Efficient Training for Neural Network Quantization
- An Adaptive and Momental Bound Method for Stochastic Learning
- Scaling Imitation Learning in Minecraft
- Deep Connectomics Networks: Neural Network Architectures Inspired by Neuronal Networks
- Mean Shift Rejection: Training Deep Neural Networks Without Minibatch Statistics or Normalization
- How many winning tickets are there in one DNN?
- On the Demystification of Knowledge Distillation: A Residual Network Perspective
- Seven Myths in Machine Learning Research
- Farkas layers: don't shift the data, fix the geometry
- On the generalization of bayesian deep nets for multi-class classification
- Modeling from Features: a Mean-field Framework for Over-parameterized Deep Neural Networks