Hilbert space embeddings and metrics on probability measures
arXiv:0907.5309
Abstract
A Hilbert space embedding for probability measures has recently been proposed, with applications including dimensionality reduction, homogeneity testing, and independence testing. This embedding represents any probability measure as a mean element in a reproducing kernel Hilbert space (RKHS). A pseudometric on the space of probability measures can be defined as the distance between distribution embeddings: we denote this as , indexed by the kernel function that defines the inner product in the RKHS. We present three theoretical properties of . First, we consider the question of determining the conditions on the kernel for which is a metric: such are denoted {\em characteristic kernels}. Unlike pseudometrics, a metric is zero only when two distributions coincide, thus ensuring the RKHS embedding maps all distributions uniquely (i.e., the embedding is injective). While previously published conditions may apply only in restricted circumstances (e.g. on compact domains), and are difficult to check, our conditions are straightforward and intuitive: bounded continuous strictly positive definite kernels are characteristic. Alternatively, if a bounded continuous kernel is translation-invariant on $\bb{R}^d$, then it is characteristic if and only if the support of its Fourier transform is the entire $\bb{R}^d$. Second, we show that there exist distinct distributions that are arbitrarily close in . Third, to understand the nature of the topology induced by , we relate to other popular metrics on probability measures, and present conditions on the kernel under which metrizes the weak topology.
48 pages
References in corpus (2)
Cited by in corpus (22)
- Equivalence of distance-based and RKHS-based statistics in hypothesis testing
- The Born Supremacy: Quantum Advantage and Training of an Ising Born Machine
- K-medoids Clustering of Data Sequences with Composite Distributions
- On the optimal estimation of probability measures in weak and strong topologies
- Quantized Compressive K-Means
- Finite sample properties of parametric MMD estimation: robustness to misspecification and dependence
- Nonparametric Detection of Geometric Structures over Networks
- Generalized Kernel Two-Sample Tests
- Correlations and Clustering in Wholesale Electricity Markets
- Generalized Similarity U: A Non-parametric Test of Association Based on Similarity
- Change Detection of Markov Kernels with Unknown Pre and Post Change Kernel
- Geometric Methods for Sampling, Optimisation, Inference and Adaptive Agents
- Principled analytic classifier for positive-unlabeled learning via weighted integral probability metric
- Nearly Consistent Finite Particle Estimates in Streaming Importance Sampling
- Model predictivity assessment: incremental test-set selection and accuracy evaluation
- A uniform kernel trick for high-dimensional two-sample problems
- Mixability of Integral Losses: a Key to Efficient Online Aggregation of Functional and Probabilistic Forecasts
- On a link between kernel mean maps and Fraunhofer diffraction, with an application to super-resolution beyond the diffraction limit
- Deep Unified Representation for Heterogeneous Recommendation
- Model-Free Change Point Detection for Mixing Processes
- Revisiting RIP guarantees for sketching operators on mixture models
- Adapting Machine Learning Diagnostic Models to New Populations Using a Small Amount of Data: Results from Clinical Neuroscience