LaProp: Separating Momentum and Adaptivity in Adam
arXiv:2002.04839
Abstract
We identity a by-far-unrecognized problem of Adam-style optimizers which results from unnecessary coupling between momentum and adaptivity. The coupling leads to instability and divergence when the momentum and adaptivity parameters are mismatched. In this work, we propose a method, Laprop, which decouples momentum and adaptivity in the Adam-style methods. We show that the decoupling leads to greater flexibility in the hyperparameters and allows for a straightforward interpolation between the signed gradient methods and the adaptive gradient methods. We experimentally show that Laprop has consistently improved speed and stability over Adam on a variety of tasks. We also bound the regret of Laprop on a convex problem and show that our bound differs from that of Adam by a key factor, which demonstrates its advantage.
References in corpus (6)
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- On the Convergence of Adam and Beyond
- Rainbow: Combining Improvements in Deep Reinforcement Learning
- Stochastic Gradient Descent Optimizes Over-parameterized Deep ReLU Networks
- Adaptive Gradient Methods with Dynamic Bound of Learning Rate
- Optimization for deep learning: theory and algorithms
Cited by in corpus (8)
- t-Soft Update of Target Network for Deep Reinforcement Learning
- Neural Networks Fail to Learn Periodic Functions and How to Fix It
- Learning Not to Learn in the Presence of Noisy Labels
- Adaptive and Multiple Time-scale Eligibility Traces for Online Deep Reinforcement Learning
- Convergent and Efficient Deep Q Network Algorithm
- Optimization Algorithm for Feedback and Feedforward Policies towards Robot Control Robust to Sensing Failures
- On the Distributional Properties of Adaptive Gradients
- Volumization as a Natural Generalization of Weight Decay