Gradient-based Hyperparameter Optimization Over Long Horizons
arXiv:2007.07869
Abstract
Gradient-based hyperparameter optimization has earned a widespread popularity in the context of few-shot meta-learning, but remains broadly impractical for tasks with long horizons (many gradient steps), due to memory scaling and gradient degradation issues. A common workaround is to learn hyperparameters online, but this introduces greediness which comes with a significant performance drop. We propose forward-mode differentiation with sharing (FDS), a simple and efficient algorithm which tackles memory scaling issues with forward-mode differentiation, and gradient degradation issues by sharing hyperparameters that are contiguous in time. We provide theoretical guarantees about the noise reduction properties of our algorithm, and demonstrate its efficiency empirically by differentiating through gradient steps of unrolled optimization. We consider large hyperparameter search ranges on CIFAR-10 where we significantly outperform greedy gradient-based alternatives, while achieving speedups compared to the state-of-the-art black-box methods. Code is available at: \url{https://github.com/polo5/FDS}
References in corpus (15)
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- Decoupled Weight Decay Regularization
- Neural Architecture Search with Reinforcement Learning
- On the difficulty of training Recurrent Neural Networks
- WaveNet: A Generative Model for Raw Audio
- Large Scale GAN Training for High Fidelity Natural Image Synthesis
- DARTS: Differentiable Architecture Search
- Learning to Reweight Examples for Robust Deep Learning
- Scalable Bayesian Optimization Using Deep Neural Networks
- Gradient-based Hyperparameter Optimization through Reversible Learning
- Understanding and Robustifying Differentiable Architecture Search
- Meta-Learning with Implicit Gradients
- Understanding and correcting pathologies in the training of learned optimizers
- Optimizing Millions of Hyperparameters by Implicit Differentiation
- Budgeted Training: Rethinking Deep Neural Network Training Under Resource Constraints