Gains in Power from Structured Two-Sample Tests of Means on Graphs
arXiv:1009.5173 · doi:10.1214/11-AOAS528
Abstract
We consider multivariate two-sample tests of means, where the location shift between the two populations is expected to be related to a known graph structure. An important application of such tests is the detection of differentially expressed genes between two patient populations, as shifts in expression levels are expected to be coherent with the structure of graphs reflecting gene properties such as biological process, molecular function, regulation, or metabolism. For a fixed graph of interest, we demonstrate that accounting for graph structure can yield more powerful tests under the assumption of smooth distribution shift on the graph. We also investigate the identification of non-homogeneous subgraphs of a given large graph, which poses both computational and multiple testing problems. The relevance and benefits of the proposed approach are illustrated on synthetic data and on breast cancer gene expression data analyzed in context of KEGG pathways.
References in corpus (4)
- A two-sample test for high-dimensional data with applications to gene-set testing
- The influence of feature selection methods on accuracy, stability and interpretability of molecular signatures
- Group Lasso with Overlaps: the Latent Group Lasso approach
- A More Powerful Two-Sample Test in High Dimensions using Random Projection
Cited by in corpus (4)
- A high-dimensional two-sample test for the mean using random subspaces
- PIMKL: Pathway Induced Multiple Kernel Learning
- A Bayesian nonparametric mixture model for selecting genes and gene subnetworks
- Bayesian Functional Analysis for Untargeted Metabolomics Data with Matching Uncertainty and Small Sample Sizes