The Break-Even Point on Optimization Trajectories of Deep Neural Networks
arXiv:2002.09572
Abstract
The early phase of training of deep neural networks is critical for their final performance. In this work, we study how the hyperparameters of stochastic gradient descent (SGD) used in the early phase of training affect the rest of the optimization trajectory. We argue for the existence of the "break-even" point on this trajectory, beyond which the curvature of the loss surface and noise in the gradient are implicitly regularized by SGD. In particular, we demonstrate on multiple classification tasks that using a large learning rate in the initial phase of training reduces the variance of the gradient, and improves the conditioning of the covariance of gradients. These effects are beneficial from the optimization perspective and become visible after the break-even point. Complementing prior work, we also show that using a low learning rate results in bad conditioning of the loss surface even for a neural network with batch normalization layers. In short, our work shows that key properties of the loss surface are strongly influenced by SGD in the early phase of training. We argue that studying the impact of the identified effects on generalization is a promising future direction.
Accepted as a spotlight at ICLR 2020. The last two authors contributed equally
References in corpus (6)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
- An Investigation into Neural Net Optimization via Hessian Eigenvalue Density
- Implicit Regularization in Deep Learning
- Measurements of Three-Level Hierarchical Structure in the Outliers in the Spectrum of Deepnet Hessians
- Negative eigenvalues of the Hessian in deep neural networks
Cited by in corpus (7)
- Accordion: Adaptive Gradient Communication via Critical Learning Regime Identification
- Learning Rate Annealing Can Provably Help Generalization, Even for Convex Problems
- Analyzing Monotonic Linear Interpolation in Neural Network Loss Landscapes
- Flatness is a False Friend
- Intraclass clustering: an implicit learning ability that regularizes DNNs
- Extrapolation for Large-batch Training in Deep Learning
- Logit Attenuating Weight Normalization