Hyperparameter optimization with approximate gradient
arXiv:1602.02355
Abstract
Most models in machine learning contain at least one hyperparameter to control for model complexity. Choosing an appropriate set of hyperparameters is both crucial in terms of model accuracy and computationally challenging. In this work we propose an algorithm for the optimization of continuous hyperparameters using inexact gradient information. An advantage of this method is that hyperparameters can be updated before model parameters have fully converged. We also give sufficient conditions for the global convergence of this method, based on regularity conditions of the involved functions and summability of errors. Finally, we validate the empirical performance of this method on the estimation of regularization constants of L2-regularized logistic regression and kernel Ridge regression. Empirical benchmarks indicate that our approach is highly competitive with respect to state of the art methods.
Fixes error in proof of Theorem 2
References in corpus (5)
- A Tutorial on Bayesian Optimization of Expensive Cost Functions, with Application to Active User Modeling and Hierarchical Reinforcement Learning
- Gradient-based Hyperparameter Optimization through Reversible Learning
- Freeze-Thaw Bayesian Optimization
- The structure of optimal parameters for image restoration problems
- Sequential Model-Based Ensemble Optimization
Cited by in corpus (34)
- AutoML: A Survey of the State-of-the-Art
- Bilevel Programming for Hyperparameter Optimization and Meta-Learning
- Stochastic Hyperparameter Optimization through Hypernetworks
- Forward and Reverse Gradient-Based Hyperparameter Optimization
- Learning the Effect of Registration Hyperparameters with HyperMorph
- PHOTONAI -- A Python API for Rapid Machine Learning Model Development
- A Self-Tuning Actor-Critic Algorithm
- Bilevel methods for image reconstruction
- Communication-Efficient Robust Federated Learning with Noisy Labels
- AutoEmb: Automated Embedding Dimensionality Search in Streaming Recommendations
- Differentiable Neural Input Search for Recommender Systems
- On Training Implicit Models
- Fast Efficient Hyperparameter Tuning for Policy Gradients
- Network Architecture Search for Domain Adaptation
- A Gradient-based Bilevel Optimization Approach for Tuning Hyperparameters in Machine Learning
- Improved Bilevel Model: Fast and Optimal Algorithm with Theoretical Guarantee
- Delta-STN: Efficient Bilevel Optimization for Neural Networks using Structured Response Jacobians
- The Differentiable Cross-Entropy Method
- Discrete Simulation Optimization for Tuning Machine Learning Method Hyperparameters
- A Linear Programming Enhanced Genetic Algorithm for Hyperparameter Tuning in Machine Learning
- Meta Approach to Data Augmentation Optimization
- Efficient Online Hyperparameter Optimization for Kernel Ridge Regression with Applications to Traffic Time Series Prediction
- MLtuner: System Support for Automatic Machine Learning Tuning
- Discretization-free Knowledge Gradient Methods for Bayesian Optimization
- Efficient Gradient Approximation Method for Constrained Bilevel Optimization
- Neural Generative Models for Global Optimization with Gradients
- Super-efficiency of automatic differentiation for functions defined as a minimum
- Meta Label Correction for Noisy Label Learning
- AReN: Assured ReLU NN Architecture for Model Predictive Control of LTI Systems
- Parameters for the best convergence of an optimization algorithm On-The-Fly
- Far-HO: A Bilevel Programming Package for Hyperparameter Optimization and Meta-Learning
- Relax and penalize: a new bilevel approach to mixed-binary hyperparameter optimization
- Ada-Segment: Automated Multi-loss Adaptation for Panoptic Segmentation
- Optimizing generalization on the train set: a novel gradient-based framework to train parameters and hyperparameters simultaneously