Simulation-Based Hypothesis Testing of High Dimensional Means Under Covariance Heterogeneity
arXiv:1406.1939 · doi:10.1111/biom.12695
Abstract
In this paper, we study the problem of testing the mean vectors of high dimensional data in both one-sample and two-sample cases. The proposed testing procedures employ maximum-type statistics and the parametric bootstrap techniques to compute the critical values. Different from the existing tests that heavily rely on the structural conditions on the unknown covariance matrices, the proposed tests allow general covariance structures of the data and therefore enjoy wide scope of applicability in practice. To enhance powers of the tests against sparse alternatives, we further propose two-step procedures with a preliminary feature screening step. Theoretical properties of the proposed tests are investigated. Through extensive numerical experiments on synthetic datasets and an human acute lymphoblastic leukemia gene expression dataset, we illustrate the performance of the new tests and how they may provide assistance on detecting disease-associated gene-sets. The proposed methods have been implemented in an R-package HDtest and are available on CRAN.
34 pages, 10 figures; Accepted for biometrics
References in corpus (7)
- On testing the significance of sets of genes
- A two-sample test for high-dimensional data with applications to gene-set testing
- Tests alternative to higher criticism for high-dimensional means under sparsity and column-wise dependence
- Higher criticism: -values and criticism
- Local independence feature screening for nonparametric and semiparametric models by marginal empirical likelihood
- A Cramér moderate deviation theorem for Hotelling's -statistic with applications to global tests
- Multiple tests of association with biological annotation metadata
Cited by in corpus (9)
- Confidence regions for entries of a large precision matrix
- Are Discoveries Spurious? Distributions of Maximum Spurious Correlations and Their Applications
- Central limit theorems for high dimensional dependent data
- High-dimensional empirical likelihood inference
- Testing the martingale difference hypothesis in high dimension
- Better-Than-Chance Classification for Signal Detection
- Edge differentially private estimation in the -model via jittering and method of moments
- An approximate randomization test for high-dimensional two-sample Behrens-Fisher problem under arbitrary covariances
- Testing for unit roots based on sample autocovariances