Discovering and Deciphering Relationships Across Disparate Data Modalities
arXiv:1609.05148 · doi:10.7554/eLife.41690
Abstract
Understanding the relationships between different properties of data, such as whether a connectome or genome has information about disease status, is becoming increasingly important in modern biological datasets. While existing approaches can test whether two properties are related, they often require unfeasibly large sample sizes in real data scenarios, and do not provide any insight into how or why the procedure reached its decision. Our approach, "Multiscale Graph Correlation" (MGC), is a dependence test that juxtaposes previously disparate data science techniques, including k-nearest neighbors, kernel methods (such as support vector machines), and multiscale analysis (such as wavelets). Other methods typically require double or triple the number samples to achieve the same statistical power as MGC in a benchmark suite including high-dimensional and nonlinear relationships - spanning polynomial (linear, quadratic, cubic), trigonometric (sinusoidal, circular, ellipsoidal, spiral), geometric (square, diamond, W-shape), and other functions, with dimensionality ranging from 1 to 1000. Moreover, MGC uniquely provides a simple and elegant characterization of the potentially complex latent geometry underlying the relationship, providing insight while maintaining computational efficiency. In several real data applications, including brain imaging and cancer genetics, MGC is the only method that can both detect the presence of a dependency and provide specific guidance for the next experiment and/or analysis to conduct.
References in corpus (12)
- Measuring and testing dependence by correlation of distances
- Brownian distance covariance
- Kernel Mean Embedding of Distributions: A Review and Beyond
- Learning Decentralized Controllers for Robot Swarms with Graph Neural Networks
- Generative Models and Model Criticism via Optimized Maximum Mean Discrepancy
- DISCO analysis: A nonparametric extension of analysis of variance
- DELTACON: A Principled Massive-Graph Similarity Function
- Large-Scale Kernel Methods for Independence Testing
- A fast algorithm for computing distance correlation
- From Distance Correlation to Multiscale Graph Correlation
- The Exact Equivalence of Distance and Kernel Methods for Hypothesis Testing
- On Quantifying Dependence: A Framework for Developing Interpretable Measures
Cited by in corpus (13)
- The Chi-Square Test of Distance Correlation
- From Distance Correlation to Multiscale Graph Correlation
- The Exact Equivalence of Distance and Kernel Methods for Hypothesis Testing
- hyppo: A Multivariate Hypothesis Testing Python Package
- Universally Consistent K-Sample Tests via Dependence Measures
- Network Dependence Testing via Diffusion Maps and Distance-Based Correlations
- Discovering the Signal Subgraph: An Iterative Screening Approach on Graphs
- Community Correlations and Testing Independence Between Binary Graphs
- Random Forests for Adaptive Nearest Neighbor Estimation of Information-Theoretic Quantities
- High-Dimensional Independence Testing via Maximum and Average Distance Correlations
- Learning Interpretable Characteristic Kernels via Decision Forests
- Efficient Bayesian Optimization using Multiscale Graph Correlation
- Valid Two-Sample Graph Testing via Optimal Transport Procrustes and Multiscale Graph Correlation with Applications in Connectomics