GradInit: Learning to Initialize Neural Networks for Stable and Efficient Training
arXiv:2102.08098
Abstract
Innovations in neural architectures have fostered significant breakthroughs in language modeling and computer vision. Unfortunately, novel architectures often result in challenging hyper-parameter choices and training instability if the network parameters are not properly initialized. A number of architecture-specific initialization schemes have been proposed, but these schemes are not always portable to new architectures. This paper presents GradInit, an automated and architecture agnostic method for initializing neural networks. GradInit is based on a simple heuristic; the norm of each network layer is adjusted so that a single step of SGD or Adam with prescribed hyperparameters results in the smallest possible loss value. This adjustment is done by introducing a scalar multiplier variable in front of each parameter block, and then optimizing these variables using a simple numerical scheme. GradInit accelerates the convergence and test performance of many convolutional architectures, both with or without skip connections, and even without normalization layers. It also improves the stability of the original Transformer architecture for machine translation, enabling training it without learning rate warmup using either Adam or SGD under a wide range of learning rates and momentum coefficients. Code is available at https://github.com/zhuchen03/gradinit.
NeurIPS 2021, fixing typos
References in corpus (8)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Improved Regularization of Convolutional Neural Networks with Cutout
- Searching for Activation Functions
- High-Performance Large-Scale Image Recognition Without Normalization
- Traditional and Heavy-Tailed Self Regularization in Neural Network Models
- Understanding the Difficulty of Training Transformers
- Characterizing signal propagation to close the performance gap in unnormalized ResNets
Cited by in corpus (7)
- NormFormer: Improved Transformer Pretraining with Extra Normalization
- PVG: Progressive Vision Graph for Vision Recognition
- Fast Certified Robust Training with Short Warmup
- The Devil is in the Detail: Simple Tricks Improve Systematic Generalization of Transformers
- Parameter Prediction for Unseen Deep Architectures
- Data-driven Weight Initialization with Sylvester Solvers
- Spending Your Winning Lottery Better After Drawing It