Preconditioning Kernel Matrices
arXiv:1602.06693
Abstract
The computational and storage complexity of kernel machines presents the primary barrier to their scaling to large, modern, datasets. A common way to tackle the scalability issue is to use the conjugate gradient algorithm, which relieves the constraints on both storage (the kernel matrix need not be stored) and computation (both stochastic gradients and parallelization can be used). Even so, conjugate gradient is not without its own issues: the conditioning of kernel matrices is often such that conjugate gradients will have poor convergence in practice. Preconditioning is a common approach to alleviating this issue. Here we propose preconditioned conjugate gradients for kernel machines, and develop a broad range of preconditioners particularly useful for kernel matrices. We describe a scalable approach to both solving kernel machines and learning their hyperparameters. We show this approach is exact in the limit of iterations and outperforms state-of-the-art approximations for a given computational budget.
References in corpus (2)
Cited by in corpus (14)
- Machine Learning Force Fields
- When Gaussian Process Meets Big Data: A Review of Scalable GPs
- Exact Gaussian Processes on a Million Data Points
- Fast Matrix Square Roots with Applications to Gaussian Processes and Bayesian Optimization
- Molecular Energy Learning Using Alternative Blackbox Matrix-Matrix Multiplication Algorithm for Exact Gaussian Process
- Fast and Accurate Gaussian Kernel Ridge Regression Using Matrix Decompositions for Preconditioning
- Bias-Free Scalable Gaussian Processes via Randomized Truncations
- Quadruply Stochastic Gaussian Processes
- Tighter Bounds on the Log Marginal Likelihood of Gaussian Process Regression Using Conjugate Gradients
- Marginalising over Stationary Kernels with Bayesian Quadrature
- A Randomized Algorithm for Preconditioner Selection
- Linear-time inference for Gaussian Processes on one dimension
- Hierarchical Inducing Point Gaussian Process for Inter-domain Observations
- Scalable Grouped Gaussian Processes via Direct Cholesky Functional Representations