Stochastic Hyperparameter Optimization through Hypernetworks
arXiv:1802.09419
Abstract
Machine learning models are often tuned by nesting optimization of model weights inside the optimization of hyperparameters. We give a method to collapse this nested optimization into joint stochastic optimization of weights and hyperparameters. Our process trains a neural network to output approximately optimal weights as a function of hyperparameters. We show that our technique converges to locally optimal weights and hyperparameters for sufficiently large hypernetworks. We compare this method to standard hyperparameter optimization strategies and demonstrate its effectiveness for tuning thousands of hyperparameters.
9 pages, 6 figures; revised figures
References in corpus (6)
- Practical Bayesian Optimization of Machine Learning Algorithms
- Weight Uncertainty in Neural Networks
- Gradient-based Hyperparameter Optimization through Reversible Learning
- SMASH: One-Shot Model Architecture Search through HyperNetworks
- Freeze-Thaw Bayesian Optimization
- Gradient-based Regularization Parameter Selection for Problems with Non-smooth Penalty Functions
Cited by in corpus (42)
- Learning to Reweight Examples for Robust Deep Learning
- A disciplined approach to neural network hyper-parameters: Part 1 -- learning rate, batch size, momentum, and weight decay
- Hypernetwork functional image representation
- Optimizing Millions of Hyperparameters by Implicit Differentiation
- Learning the Effect of Registration Hyperparameters with HyperMorph
- Principled Weight Initialization for Hypernetworks
- Personalized Federated Learning using Hypernetworks
- Hyperparameter Ensembles for Robustness and Uncertainty Quantification
- Learning the Pareto Front with Hypernetworks
- Communication-Efficient Robust Federated Learning with Noisy Labels
- HyperGAN: A Generative Model for Diverse, Performant Neural Networks
- Regularization Learning Networks: Deep Learning for Tabular Datasets
- Robust Federated Learning Through Representation Matching and Adaptive Hyper-parameters
- Beyond backpropagation: bilevel optimization through implicit differentiation and equilibrium propagation
- On the Modularity of Hypernetworks
- A Generic First-Order Algorithmic Framework for Bi-Level Programming Beyond Lower-Level Singleton
- A Gradient-based Bilevel Optimization Approach for Tuning Hyperparameters in Machine Learning
- Regularization-Agnostic Compressed Sensing MRI Reconstruction with Hypernetworks
- Generating Neural Networks with Neural Networks
- Improved Bilevel Model: Fast and Optimal Algorithm with Theoretical Guarantee
- Delta-STN: Efficient Bilevel Optimization for Neural Networks using Structured Response Jacobians
- Meta-MgNet: Meta Multigrid Networks for Solving Parameterized Partial Differential Equations
- Learning to Impute: A General Framework for Semi-supervised Learning
- On Infinite-Width Hypernetworks
- Combiner and HyperCombiner Networks: Rules to Combine Multimodality MR Images for Prostate Cancer Localisation
- Hidden Incentives for Auto-Induced Distributional Shift
- A Linear Programming Enhanced Genetic Algorithm for Hyperparameter Tuning in Machine Learning
- EvoGrad: Efficient Gradient-Based Meta-Learning and Hyperparameter Optimization
- OnlineAugment: Online Data Augmentation with Less Domain Knowledge
- Towards Evaluating the Robustness of Neural Networks Learned by Transduction
- Teaching with Commentaries
- A Generative Model for Sampling High-Performance and Diverse Weights for Neural Networks
- Stability and Generalization of Bilevel Programming in Hyperparameter Optimization
- Towards Adversarial Robustness via Transductive Learning
- Online hyperparameter optimization by real-time recurrent learning
- MetaInv-Net: Meta Inversion Network for Sparse View CT Image Reconstruction
- Stabilizing Bi-Level Hyperparameter Optimization using Moreau-Yosida Regularization
- Cost-Efficient Online Hyperparameter Optimization
- Meta-Learning to Improve Pre-Training
- Meta Internal Learning
- HALO: Learning to Prune Neural Networks with Shrinkage
- Complex Momentum for Optimization in Games