A high-dimensional two-sample test for the mean using random subspaces
arXiv:1304.4564 · doi:10.1016/j.csda.2013.12.003
Abstract
A common problem in genetics is that of testing whether a set of highly dependent gene expressions differ between two populations, typically in a high-dimensional setting where the data dimension is larger than the sample size. Most high-dimensional tests for the equality of two mean vectors rely on naive diagonal or trace estimators of the covariance matrix, ignoring dependencies between variables. A test recently proposed by Lopes et al. (2012) implicitly incorporates dependencies by using random pseudo-projections to a lower-dimensional space. Their test offers higher power when the variables are dependent, but lacks desirable invariance properties and relies on asymptotic p-values that are too conservative. We illustrate how a permutation approach can be used to obtain p-values for the Lopes et al. test and how modifying the test using random subspaces leads to a test statistic that is invariant under linear transformations of the marginal distributions. The resulting test does not rely on assumptions about normality or the structure of the covariance matrix. We show by simulation that the new test has higher power than competing tests in realistic settings motivated by microarray gene expression data. We also discuss the computational aspects of high-dimensional permutation tests and provide an efficient R implementation of the proposed test.
References in corpus (7)
- On testing the significance of sets of genes
- A two-sample test for high-dimensional data with applications to gene-set testing
- Two sample tests for high-dimensional covariance matrices
- Random-set methods identify distinct aspects of the enrichment signal in gene-set analysis
- More power via graph-structured tests for differential expression of gene networks
- A statistical framework for testing functional categories in microarray data
- Direction-Projection-Permutation for High Dimensional Hypothesis Tests
Cited by in corpus (4)
- A review of 20 years of naive tests of significance for high-dimensional mean vectors and covariance matrices
- On split sample and randomized confidence intervals for binomial proportions
- A Neighborhood-Assisted Hotelling's Test for High-Dimensional Means
- Diagonal Likelihood Ratio Test for Equality of Mean Vectors in High-Dimensional Data