The Exact Equivalence of Distance and Kernel Methods for Hypothesis Testing
arXiv:1806.05514 · doi:10.1007/s10182-020-00378-1
Abstract
Distance-based tests, also called "energy statistics", are leading methods for two-sample and independence tests from the statistics community. Kernel-based tests, developed from "kernel mean embeddings", are leading methods for two-sample and independence tests from the machine learning community. A fixed-point transformation was previously proposed to connect the distance methods and kernel methods for the population statistics. In this paper, we propose a new bijective transformation between metrics and kernels. It simplifies the fixed-point transformation, inherits similar theoretical properties, allows distance methods to be exactly the same as kernel methods for sample statistics and p-value, and better preserves the data structure upon transformation. Our results further advance the understanding in distance and kernel-based tests, streamline the code base for implementing these tests, and enable a rich literature of distance-based and kernel-based methodologies to directly communicate with each other.
24 pages main + 7 pages appendix, 3 figures
References in corpus (8)
- Measuring and testing dependence by correlation of distances
- Brownian distance covariance
- DISCO analysis: A nonparametric extension of analysis of variance
- The Chi-Square Test of Distance Correlation
- A fast algorithm for computing distance correlation
- From Distance Correlation to Multiscale Graph Correlation
- Community Correlations and Testing Independence Between Binary Graphs
- High-Dimensional Independence Testing via Maximum and Average Distance Correlations
Cited by in corpus (13)
- The Chi-Square Test of Distance Correlation
- Review of end-to-end speech synthesis technology based on deep learning
- hyppo: A Multivariate Hypothesis Testing Python Package
- Universally Consistent K-Sample Tests via Dependence Measures
- Synergistic Graph Fusion via Encoder Embedding
- Correcting a Nonparametric Two-sample Graph Hypothesis Test for Graphs with Different Numbers of Vertices with Applications to Connectomics
- Encoder Embedding for General Graph and Node Classification
- Discovering the Signal Subgraph: An Iterative Screening Approach on Graphs
- Learning Interpretable Characteristic Kernels via Decision Forests
- High-Dimensional Independence Testing via Maximum and Average Distance Correlations
- Fast and Scalable Multi-Kernel Encoder Classifier
- Multiscale Comparative Connectomics
- Private measurement of nonlinear correlations between data hosted across multiple parties