Learning Gradient Descent: Better Generalization and Longer Horizons
arXiv:1703.03633
Abstract
Training deep neural networks is a highly nontrivial task, involving carefully selecting appropriate training algorithms, scheduling step sizes and tuning other hyperparameters. Trying different combinations can be quite labor-intensive and time consuming. Recently, researchers have tried to use deep learning algorithms to exploit the landscape of the loss function of the training problem of interest, and learn how to optimize over it in an automatic way. In this paper, we propose a new learning-to-learn model and some useful and practical tricks. Our optimizer outperforms generic, hand-crafted optimization algorithms and state-of-the-art learning-to-learn optimizers by DeepMind in many tasks. We demonstrate the effectiveness of our algorithms on a number of tasks, including deep MLPs, CNNs, and simple LSTMs.
Accepted to ICML 2017, 9 pages, 9 figures, 4 tables
References in corpus (3)
Cited by in corpus (17)
- Learning to Optimize: A Primer and A Benchmark
- Deep residual detection of radio frequency interference for FAST
- Understanding Short-Horizon Bias in Stochastic Meta-Optimization
- Meta Continual Learning
- Learning to Optimize in Model Predictive Control
- Tasks, stability, architecture, and compute: Training more effective learned optimizers, and using them to train themselves
- Using a thousand optimization tasks to learn hyperparameter search strategies
- Training Stronger Baselines for Learning to Optimize
- Guarantees for Tuning the Step Size using a Learning-to-Learn Approach
- Learning to be Global Optimizer
- ROAM: Recurrently Optimizing Tracking Model
- Reverse engineering learned optimizers reveals known and novel mechanisms
- Learning to Optimize with Dynamic Mode Decomposition
- Gradients are Not All You Need
- Adaptive Hierarchical Hyper-gradient Descent
- ModelPred: A Framework for Predicting Trained Model from Training Data
- Bootstrapped Meta-Learning