Bayesian Nonparametric Kernel-Learning
arXiv:1506.08776
Abstract
Kernel methods are ubiquitous tools in machine learning. However, there is often little reason for the common practice of selecting a kernel a priori. Even if a universal approximating kernel is selected, the quality of the finite sample estimator may be greatly affected by the choice of kernel. Furthermore, when directly applying kernel methods, one typically needs to compute a Gram matrix of pairwise kernel evaluations to work with a dataset of instances. The computation of this Gram matrix precludes the direct application of kernel methods on large datasets, and makes kernel learning especially difficult. In this paper we introduce Bayesian nonparmetric kernel-learning (BaNK), a generic, data-driven framework for scalable learning of kernels. BaNK places a nonparametric prior on the spectral distribution of random frequencies allowing it to both learn kernels and scale to large datasets. We show that this framework can be used for large scale regression and classification tasks. Furthermore, we show that BaNK outperforms several other scalable approaches for kernel learning on a variety of real world datasets.
Cited by in corpus (16)
- Synthesizing Tabular Data using Generative Adversarial Networks
- Linear Multiple Low-Rank Kernel Based Stationary Gaussian Processes Regression for Time Series
- Random Features for Kernel Approximation: A Survey on Algorithms, Theory, and Beyond
- On Sampling Random Features From Empirical Leverage Scores: Implementation and Theoretical Guarantees
- A Statistical Approach to Surface Metrology for 3D-Printed Stainless Steel
- On Learning the Transformer Kernel
- Generalization Guarantees for Sparse Kernel Approximation with Entropic Optimal Features
- Pseudo-Bayesian Learning with Kernel Fourier Transform as Prior
- Latent variable modeling with random features
- Marginalising over Stationary Kernels with Bayesian Quadrature
- ORCCA: Optimal Randomized Canonical Correlation Analysis
- Not-So-Random Features
- Efficient and Adaptive Kernelization for Nonlinear Max-margin Multi-view Learning
- Gaussian Processes with Skewed Laplace Spectral Mixture Kernels for Long-term Forecasting
- Learning Causally-Generated Stationary Time Series
- Learning Landmark-Based Ensembles with Random Fourier Features and Gradient Boosting