Stochastic gradient descent with random learning rate
arXiv:2003.06926
Abstract
We propose to optimize neural networks with a uniformly-distributed random learning rate. The associated stochastic gradient descent algorithm can be approximated by continuous stochastic equations and analyzed within the Fokker-Planck formalism. In the small learning rate regime, the training process is characterized by an effective temperature which depends on the average learning rate, the mini-batch size and the momentum of the optimization algorithm. By comparing the random learning rate protocol with cyclic and constant protocols, we suggest that the random choice is generically the best strategy in the small learning rate regime, yielding better regularization without extra computational cost. We provide supporting evidence through experiments on both shallow, fully-connected and deep, convolutional neural networks for image classification on the MNIST and CIFAR10 datasets.
improved theoretical and experimental analysis
References in corpus (9)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Machine learning and the physical sciences
- Opening the Black Box of Deep Neural Networks via Information
- Adding noise to the input of a model trained with a regularized objective
- The Implicit Regularization of Stochastic Gradient Flow for Least Squares
- Maximum Entropy approach to multivariate time series randomization
- Partial local entropy and anisotropy in deep weight spaces
- Correspondence between temporal correlations in time series, inverse problems, and the Spherical Model
- Training Deep Neural Networks by optimizing over nonlocal paths in hyperparameter space