The large learning rate phase of deep learning: the catapult mechanism
arXiv:2003.02218
Abstract
The choice of initial learning rate can have a profound effect on the performance of deep networks. We present a class of neural networks with solvable training dynamics, and confirm their predictions empirically in practical deep learning settings. The networks exhibit sharply distinct behaviors at small and large learning rates. The two regimes are separated by a phase transition. In the small learning rate phase, training can be understood using the existing theory of infinitely wide neural networks. At large learning rates the model captures qualitatively distinct phenomena, including the convergence of gradient descent dynamics to flatter minima. One key prediction of our model is a narrow range of large, stable learning rates. We find good agreement between our model's predictions and training dynamics in realistic deep learning settings. Furthermore, we find that the optimal performance in such settings is often found in the large learning rate phase. We believe our results shed light on characteristics of models trained at different learning rates. In particular, they fill a gap between existing wide neural network theory, and the nonlinear, large learning rate, training dynamics relevant to practice.
25 pages, 19 figures
References in corpus (2)
Cited by in corpus (15)
- Scaling Laws for Autoregressive Generative Modeling
- Understanding the Role of Training Regimes in Continual Learning
- On the Origin of Implicit Regularization in Stochastic Gradient Descent
- Finite Versus Infinite Neural Networks: an Empirical Study
- Deep learning versus kernel learning: an empirical study of loss landscape geometry and the time evolution of the Neural Tangent Kernel
- How to decay your learning rate
- The Surprising Simplicity of the Early-Time Learning Dynamics of Neural Networks
- Differentiable Physics: A Position Piece
- Learning Rate Annealing Can Provably Help Generalization, Even for Convex Problems
- Traces of Class/Cross-Class Structure Pervade Deep Learning Spectra
- Analyzing Monotonic Linear Interpolation in Neural Network Loss Landscapes
- A Loss Curvature Perspective on Training Instability in Deep Learning
- Implicit bias of deep linear networks in the large learning rate phase
- Transformed CNNs: recasting pre-trained convolutional layers with self-attention
- Multilayer Lookahead: a Nested Version of Lookahead