Scalable Gradient-Based Tuning of Continuous Regularization Hyperparameters
arXiv:1511.06727
Abstract
Hyperparameter selection generally relies on running multiple full training trials, with selection based on validation set performance. We propose a gradient-based approach for locally adjusting hyperparameters during training of the model. Hyperparameters are adjusted so as to make the model parameter gradients, and hence updates, more advantageous for the validation cost. We explore the approach for tuning regularization hyperparameters and find that in experiments on MNIST, SVHN and CIFAR-10, the resulting regularization levels are within the optimal regions. The additional computational cost depends on how frequently the hyperparameters are trained, but the tested scheme adds only 30% computational overhead regardless of the model size. Since the method is significantly less computationally demanding compared to similar gradient-based approaches to hyperparameter optimization, and consistently finds good hyperparameter values, it can be a useful tool for training neural network models.
9 pages, 7 figures. Accepted at ICML 2016
References in corpus (5)
Cited by in corpus (39)
- An Ensemble of Epoch-wise Empirical Bayes for Few-shot Learning
- Generalized Inner Loop Meta-Learning
- Coresets via Bilevel Optimization for Continual Learning and Streaming
- Efficient Hyperparameter Optimization in Deep Learning Using a Variable Length Genetic Algorithm
- Stochastic Hyperparameter Optimization through Hypernetworks
- Optimizing Millions of Hyperparameters by Implicit Differentiation
- Learning the Effect of Registration Hyperparameters with HyperMorph
- DeepEMD: Differentiable Earth Mover's Distance for Few-Shot Learning
- Penalty Method for Inversion-Free Deep Bilevel Optimization
- Probabilistic Dual Network Architecture Search on Graphs
- Gradient Descent: The Ultimate Optimizer
- Not All Unlabeled Data are Equal: Learning to Weight Data in Semi-supervised Learning
- Truncated Back-propagation for Bilevel Optimization
- Flexible Dataset Distillation: Learn Labels Instead of Images
- Automatic Mixed-Precision Quantization Search of BERT
- On Training Implicit Models
- Fast Efficient Hyperparameter Tuning for Policy Gradients
- Network Architecture Search for Domain Adaptation
- Improved Bilevel Model: Fast and Optimal Algorithm with Theoretical Guarantee
- Delta-STN: Efficient Bilevel Optimization for Neural Networks using Structured Response Jacobians
- MetaMixUp: Learning Adaptive Interpolation Policy of MixUp with Meta-Learning
- EvoGrad: Efficient Gradient-Based Meta-Learning and Hyperparameter Optimization
- The Nonlinearity Coefficient - A Practical Guide to Neural Architecture Design
- Auxiliary Learning by Implicit Differentiation
- Online Hyperparameter Meta-Learning with Hypergradient Distillation
- Novel Suboptimal approaches for Hyperparameter Tuning of Deep Neural Network [under the shelf of Optical Communication]
- A contrastive rule for meta-learning
- MetaInv-Net: Meta Inversion Network for Sparse View CT Image Reconstruction
- Stability and Generalization of Bilevel Programming in Hyperparameter Optimization
- Online hyperparameter optimization by real-time recurrent learning
- One-step differentiation of iterative algorithms
- Hyperparameter Optimization in Neural Networks via Structured Sparse Recovery
- Graduated Optimization of Black-Box Functions
- Gradient-based Hyperparameter Optimization Over Long Horizons
- Efficient Hyperparameter Tuning with Dynamic Accuracy Derivative-Free Optimization
- Cost-Efficient Online Hyperparameter Optimization
- In-Loop Meta-Learning with Gradient-Alignment Reward
- Task-Driven Data Verification via Gradient Descent
- Automation for Interpretable Machine Learning Through a Comparison of Loss Functions to Regularisers