Probabilistic Line Searches for Stochastic Optimization
arXiv:1502.02846
Abstract
In deterministic optimization, line searches are a standard tool ensuring stability and efficiency. Where only stochastic gradients are available, no direct equivalent has so far been formulated, because uncertain gradients do not allow for a strict sequence of decisions collapsing the search space. We construct a probabilistic line search by combining the structure of existing deterministic methods with notions from Bayesian optimization. Our method retains a Gaussian process surrogate of the univariate optimization objective, and uses a probabilistic belief over the Wolfe conditions to monitor the descent. The algorithm has very low computational cost, and no user-controlled parameters. Experiments show that it effectively removes the need to define a learning rate for stochastic gradient descent.
12 pages, including supplements
References in corpus (2)
Cited by in corpus (7)
- Dissecting Adam: The Sign, Magnitude and Variance of Stochastic Gradients
- AutoLRS: Automatic Learning-Rate Schedule by Bayesian Optimization on the Fly
- LOSSGRAD: automatic learning rate in gradient descent
- Optimality Criteria for Probabilistic Numerical Methods
- Faster Convergence for Transformer Fine-tuning with Line Search Methods
- Improving Line Search Methods for Large Scale Neural Network Training
- Black Box Probabilistic Numerics