The Parallel Knowledge Gradient Method for Batch Bayesian Optimization
arXiv:1606.04414
Abstract
In many applications of black-box optimization, one can evaluate multiple points simultaneously, e.g. when evaluating the performances of several different neural network architectures in a parallel computing environment. In this paper, we develop a novel batch Bayesian optimization algorithm --- the parallel knowledge gradient method. By construction, this method provides the one-step Bayes-optimal batch of points to sample. We provide an efficient strategy for computing this Bayes-optimal batch of points, and we demonstrate that the parallel knowledge gradient method finds global optima significantly faster than previous batch Bayesian optimization algorithms on both synthetic test functions and when tuning hyperparameters of practical machine learning algorithms, especially when function evaluations are noisy.
Minor edits and typo fixes. Please cite "J. Wu and P. Frazier. The parallel knowledge gradient method for batch bayesian optimization. In Advances In Neural Information Processing Systems, pp. 3126-3134. 2016"
Cited by in corpus (25)
- A Tutorial on Bayesian Optimization
- Bayesian Optimization is Superior to Random Search for Machine Learning Hyperparameter Tuning: Analysis of the Black-Box Optimization Challenge 2020
- Differentiable Expected Hypervolume Improvement for Parallel Multi-Objective Bayesian Optimization
- pySOT and POAP: An event-driven asynchronous framework for surrogate optimization
- Is novelty predictable?
- Interpolating Detailed Simulations of Kilonovae: Adaptive Learning and Parameter Inference Applications
- Bayesian Optimisation over Multiple Continuous and Categorical Inputs
- Batch simulations and uncertainty quantification in Gaussian process surrogate approximate Bayesian computation
- GIBBON: General-purpose Information-Based Bayesian OptimisatioN
- Towards Assessing the Impact of Bayesian Optimization's Own Hyperparameters
- Sampling Acquisition Functions for Batch Bayesian Optimization
- High-Dimensional Experimental Design and Kernel Bandits
- Bayesian Optimization of Risk Measures
- Knowledge Gradient for Selection with Covariates: Consistency and Computation
- Scalable Thompson Sampling using Sparse Gaussian Process Models
- BINOCULARS for Efficient, Nonmyopic Sequential Experimental Design
- Bayesian Optimization for Iterative Learning
- Practical Batch Bayesian Optimization for Less Expensive Functions
- Efficient Batch Black-box Optimization with Deterministic Regret Bounds
- Adversarial Likelihood-Free Inference on Black-Box Generator
- Uncertainty Quantification for Bayesian Optimization
- Efficient Phase Diagram Sampling by Active Learning
- Automatic Calibration of Dynamic and Heterogeneous Parameters in Agent-based Model
- A tree-based radial basis function method for noisy parallel surrogate optimization
- HyperZero: A Customized End-to-End Auto-Tuning System for Recommendation with Hourly Feedback