The Random Forest Kernel and other kernels for big data from random partitions
arXiv:1402.4293
Abstract
We present Random Partition Kernels, a new class of kernels derived by demonstrating a natural connection between random partitions of objects and kernels between those objects. We show how the construction can be used to create kernels from methods that would not normally be viewed as random partitions, such as Random Forest. To demonstrate the potential of this method, we propose two new kernels, the Random Forest Kernel and the Fast Cluster Kernel, and show that these kernels consistently outperform standard kernels on problems involving real-world datasets. Finally, we show how the form of these kernels lend themselves to a natural approximation that is appropriate for certain big data problems, allowing inference in methods such as Gaussian Processes, Support Vector Machines and Kernel PCA.
Cited by in corpus (10)
- What do you Mean? The Role of the Mean Function in Bayesian Optimisation
- Boulevard: Regularized Stochastic Gradient Boosted Trees and Their Limiting Distribution
- Robust Hypothesis Test for Nonlinear Effect with Gaussian Processes
- Learning Interpretable Characteristic Kernels via Decision Forests
- A Novel Random Forest Dissimilarity Measure for Multi-View Learning
- Random forests and kernel methods
- TREX: Tree-Ensemble Representer-Point Explanations
- Critères de qualité d'un classifieur généraliste
- Geodesic Learning via Unsupervised Decision Forests
- A Framework for an Assessment of the Kernel-target Alignment in Tree Ensemble Kernel Learning