Thoughts on Massively Scalable Gaussian Processes
arXiv:1511.01870
Abstract
We introduce a framework and early results for massively scalable Gaussian processes (MSGP), significantly extending the KISS-GP approach of Wilson and Nickisch (2015). The MSGP framework enables the use of Gaussian processes (GPs) on billions of datapoints, without requiring distributed inference, or severe assumptions. In particular, MSGP reduces the standard complexity of GP learning and inference to , and the standard complexity per test point prediction to . MSGP involves 1) decomposing covariance matrices as Kronecker products of Toeplitz matrices approximated by circulant matrices. This multi-level circulant approximation allows one to unify the orthogonal computational benefits of fast Kronecker and Toeplitz approaches, and is significantly faster than either approach in isolation; 2) local kernel interpolation and inducing points to allow for arbitrarily located data inputs, and test time predictions; 3) exploiting block-Toeplitz Toeplitz-block structure (BTTB), which enables fast inference and learning when multidimensional Kronecker structure is not present; and 4) projections of the input space to flexibly model correlated inputs and high dimensional data. The ability to handle many () inducing points allows for near-exact accuracy and large scale kernel learning.
25 pages, 9 figures
Cited by in corpus (31)
- Fast and scalable Gaussian process modeling with applications to astronomical time series
- GPyTorch: Blackbox Matrix-Matrix Gaussian Process Inference with GPU Acceleration
- A quantum linear system algorithm for dense matrices
- Adversarial Examples, Uncertainty, and Transfer Testing Robustness in Gaussian Process Hybrid Deep Networks
- When Gaussian Process Meets Big Data: A Review of Scalable GPs
- Local approximate Gaussian process regression for data-driven constitutive laws: Development and comparison with neural networks
- GPLaSDI: Gaussian Process-based Interpretable Latent Space Dynamics Identification through Deep Autoencoder
- Recurrent Attentive Neural Process for Sequential Data
- Algorithmic Linearly Constrained Gaussian Processes
- Quantum algorithms for scientific computing
- Deep Kernel Learning
- Constant-Time Predictive Distributions for Gaussian Processes
- State Space Gaussian Processes with Non-Gaussian Likelihood
- Spectral Convergence of Graph Laplacian and Heat Kernel Reconstruction in from Random Samples
- Multitask methods for predicting molecular properties from heterogeneous data
- Lifelong Bayesian Optimization
- Fast increased fidelity approximate Gibbs samplers for Bayesian Gaussian process regression
- Distilled Thompson Sampling: Practical and Efficient Thompson Sampling via Imitation Learning
- SKIing on Simplices: Kernel Interpolation on the Permutohedral Lattice for Scalable Gaussian Processes
- Fast Kronecker Matrix-Matrix Multiplication on GPUs
- Diverse Video Generation using a Gaussian Process Trigger
- Understanding Uncertainty in Bayesian Deep Learning
- UNITE: Uncertainty-based Health Risk Prediction Leveraging Multi-sourced Data
- Scalable Variational Gaussian Processes via Harmonic Kernel Decomposition
- Large-scale magnetic field maps using structured kernel interpolation for Gaussian process regression
- Scalable Gaussian Process Regression for Kernels with a Non-Stationary Phase
- Bayesian Inference in High-Dimensional Time-Serieswith the Orthogonal Stochastic Linear Mixing Model
- Hierarchical Inducing Point Gaussian Process for Inter-domain Observations
- Faster Kernel Interpolation for Gaussian Processes
- Scalable Bayesian Optimization with Sparse Gaussian Process Models
- Sparse within Sparse Gaussian Processes using Neighbor Information