Input Warping for Bayesian Optimization of Non-stationary Functions
arXiv:1402.0929
Abstract
Bayesian optimization has proven to be a highly effective methodology for the global optimization of unknown, expensive and multimodal functions. The ability to accurately model distributions over functions is critical to the effectiveness of Bayesian optimization. Although Gaussian processes provide a flexible prior over functions which can be queried efficiently, there are various classes of functions that remain difficult to model. One of the most frequently occurring of these is the class of non-stationary functions. The optimization of the hyperparameters of machine learning algorithms is a problem domain in which parameters are often manually transformed a priori, for example by optimizing in "log-space," to mitigate the effects of spatially-varying length scale. We develop a methodology for automatically learning a wide family of bijective transformations or warpings of the input space using the Beta cumulative distribution function. We further extend the warping framework to multi-task Bayesian optimization so that multiple tasks can be warped into a jointly stationary space. On a set of challenging benchmark optimization tasks, we observe that the inclusion of warping greatly improves on the state-of-the-art, producing better results faster and more reliably.
References in corpus (5)
- Improving neural networks by preventing co-adaptation of feature detectors
- Practical Bayesian Optimization of Machine Learning Algorithms
- Slice sampling covariance hyperparameters of latent Gaussian models
- Portfolio Allocation for Bayesian Optimization
- Anomaly Detection and Removal Using Non-Stationary Gaussian Processes
Cited by in corpus (32)
- Scalable Bayesian Optimization Using Deep Neural Networks
- Generative Moment Matching Networks
- Practical heteroskedastic Gaussian process modeling for large simulation experiments
- An Ensemble of Epoch-wise Empirical Bayes for Few-shot Learning
- Learning to Optimize: A Primer and A Benchmark
- Bayesian Optimization with Adaptive Kernels for Robot Control
- Revisiting Bayesian Optimization in the light of the COCO benchmark
- A Stratified Analysis of Bayesian Optimization Methods
- Scalable Constrained Bayesian Optimization
- Bag of Baselines for Multi-objective Joint Neural Architecture Search and Hyperparameter Optimization
- Multi-Objective Bayesian Optimization over High-Dimensional Search Spaces
- Multi-level CNN for lung nodule classification with Gaussian Process assisted hyperparameter optimization
- lgpr: An interpretable nonparametric method for inferring covariate effects from longitudinal data
- Fast Efficient Hyperparameter Tuning for Policy Gradients
- Single Gaussian Process Method for Arbitrary Tokamak Regimes with a Statistical Analysis
- Better call Surrogates: A hybrid Evolutionary Algorithm for Hyperparameter optimization
- Neural Network Architecture Optimization through Submodularity and Supermodularity
- How Bayesian Should Bayesian Optimisation Be?
- Differentially Private Gaussian Processes
- Ordinal Bayesian Optimisation
- Compositional uncertainty in deep Gaussian processes
- Sparse Spectrum Warped Input Measures for Nonstationary Kernel Learning
- Non-smooth Bayesian Optimization in Tuning Problems
- No-Regret Bayesian Optimization with Unknown Hyperparameters
- Monotonic Gaussian Process Flow
- BORE: Bayesian Optimization by Density-Ratio Estimation
- Gaussian Process Bandit Optimization of the Thermodynamic Variational Objective
- State-space deep Gaussian processes with applications
- Hyperparameter Optimization in Neural Networks via Structured Sparse Recovery
- Neural fidelity warping for efficient robot morphology design
- Reducing The Search Space For Hyperparameter Optimization Using Group Sparsity
- Warped Input Gaussian Processes for Time Series Forecasting