Fast Bayesian Optimization of Machine Learning Hyperparameters on Large Datasets
arXiv:1605.07079
Abstract
Bayesian optimization has become a successful tool for hyperparameter optimization of machine learning algorithms, such as support vector machines or deep neural networks. Despite its success, for large datasets, training and validating a single configuration often takes hours, days, or even weeks, which limits the achievable performance. To accelerate hyperparameter optimization, we propose a generative model for the validation error as a function of training set size, which is learned during the optimization process and allows exploration of preliminary configurations on small subsets, by extrapolating to the full dataset. We construct a Bayesian optimization procedure, dubbed Fabolas, which models loss and training time as a function of dataset size and automatically trades off high information gain about the global optimum against computational cost. Experiments optimizing support vector machines and deep neural networks show that Fabolas often finds high-quality solutions 10 to 100 times faster than other state-of-the-art Bayesian optimization methods or the recently proposed bandit strategy Hyperband.
References in corpus (9)
- Practical Bayesian Optimization of Machine Learning Algorithms
- A Tutorial on Bayesian Optimization of Expensive Cost Functions, with Application to Active User Modeling and Hierarchical Reinforcement Learning
- Expectation Propagation for approximate Bayesian inference
- Weight Uncertainty in Neural Networks
- OpenML: networked science in machine learning
- Scalable Bayesian Optimization Using Deep Neural Networks
- Predictive Entropy Search for Efficient Global Optimization of Black-box Functions
- Freeze-Thaw Bayesian Optimization
- Automated Machine Learning on Big Data using Stochastic Algorithm Tuning
Cited by in corpus (29)
- Deep learning with convolutional neural networks for EEG decoding and visualization
- A Downsampled Variant of ImageNet as an Alternative to the CIFAR datasets
- A Tutorial on Bayesian Optimization
- Neural Architecture Search with Bayesian Optimisation and Optimal Transport
- Understanding the effect of hyperparameter optimization on machine learning models for structure design problems
- Automated Reinforcement Learning (AutoRL): A Survey and Open Problems
- Towards Green Automated Machine Learning: Status Quo and Future Directions
- Learning Curves for Decision Making in Supervised Machine Learning: A Survey
- Cost-aware Bayesian Optimization
- Efficient Automatic CASH via Rising Bandits
- Robust Federated Learning Through Representation Matching and Adaptive Hyper-parameters
- HARK Side of Deep Learning -- From Grad Student Descent to Automated Machine Learning
- Efficient Visual Recognition with Deep Neural Networks: A Survey on Recent Advances and New Directions
- Hyperparameter Learning via Distributional Transfer
- Identifiability and physical interpretability of hybrid, gray-box models -- a case study
- Speedy Performance Estimation for Neural Architecture Search
- Output Space Entropy Search Framework for Multi-Objective Bayesian Optimization
- Fast Efficient Hyperparameter Tuning for Policy Gradients
- Multi-objective Bayesian optimisation with preferences over objectives
- Lifelong Bayesian Optimization
- Efficient Online Hyperparameter Optimization for Kernel Ridge Regression with Applications to Traffic Time Series Prediction
- Rethinking Performance Estimation in Neural Architecture Search
- Capacity allocation analysis of neural networks: A tool for principled architecture design
- Weighting Is Worth the Wait: Bayesian Optimization with Importance Sampling
- Reusing Trained Layers of Convolutional Neural Networks to Shorten Hyperparameters Tuning Time
- Noisy Blackbox Optimization with Multi-Fidelity Queries: A Tree Search Approach
- A Nonmyopic Approach to Cost-Constrained Bayesian Optimization
- Estimating Shape Parameters of Piecewise Linear-Quadratic Problems
- Multifidelity Bayesian Optimization for Binomial Output