Variants of RMSProp and Adagrad with Logarithmic Regret Bounds
arXiv:1706.05507
Abstract
Adaptive gradient methods have become recently very popular, in particular as they have been shown to be useful in the training of deep neural networks. In this paper we have analyzed RMSProp, originally proposed for the training of deep neural networks, in the context of online convex optimization and show -type regret bounds. Moreover, we propose two variants SC-Adagrad and SC-RMSProp for which we show logarithmic regret bounds for strongly convex functions. Finally, we demonstrate in the experiments that these new variants outperform other adaptive gradient techniques or stochastic gradient descent in the optimization of strongly convex functions as well as in training of deep neural networks.
ICML 2017, 16 pages, 23 figures
References in corpus (2)
Cited by in corpus (9)
- Adam revisited: a weighted past gradients perspective
- A Selective Review on Statistical Methods for Massive Data Computation: Distributed Computing, Subsampling, and Minibatch Techniques
- Gravity Optimizer: a Kinematic Approach on Optimization in Deep Learning
- SAdam: A Variant of Adam for Strongly Convex Functions
- BPGrad: Towards Global Optimality in Deep Learning via Branch and Pruning
- Gradient-only line searches to automatically determine learning rates for a variety of stochastic training algorithms
- Per-pixel Classification Rebar Exposures in Bridge Eye-inspection
- On Generalization of Adaptive Methods for Over-parameterized Linear Regression
- Binary Search and First Order Gradient Based Method for Stochastic Optimization