Are Discoveries Spurious? Distributions of Maximum Spurious Correlations and Their Applications
arXiv:1502.04237 · doi:10.1214/17-AOS1575
Abstract
Over the last two decades, many exciting variable selection methods have been developed for finding a small group of covariates that are associated with the response from a large pool. Can the discoveries from these data mining approaches be spurious due to high dimensionality and limited sample size? Can our fundamental assumptions about the exogeneity of the covariates needed for such variable selection be validated with the data? To answer these questions, we need to derive the distributions of the maximum spurious correlations given a certain number of predictors, namely, the distribution of the correlation of a response variable with the best linear combinations of covariates , even when and are independent. When the covariance matrix of possesses the restricted eigenvalue property, we derive such distributions for both a finite and a diverging , using Gaussian approximation and empirical process techniques. However, such a distribution depends on the unknown covariance matrix of . Hence, we use the multiplier bootstrap procedure to approximate the unknown distributions and establish the consistency of such a simple bootstrap approach. The results are further extended to the situation where the residuals are from regularized fits. Our approach is then used to construct the upper confidence limit for the maximum spurious correlation and to test the exogeneity of the covariates. The former provides a baseline for guarding against false discoveries and the latter tests whether our fundamental assumptions for high-dimensional model selection are statistically valid. Our techniques and results are illustrated with both numerical examples and real data analysis.
References in corpus (9)
- Nearly unbiased variable selection under minimax concave penalty
- Challenges of Big Data Analysis
- One-step sparse estimates in nonconcave penalized likelihood models
- Distributions of Angles in Random Packing on Spheres
- Simulation-Based Hypothesis Testing of High Dimensional Means Under Covariance Heterogeneity
- A Selective Overview of Variable Selection in High Dimensional Feature Space (Invited Review Article)
- Non-Concave Penalized Likelihood with NP-Dimensionality
- Necessary and sufficient conditions for the asymptotic distributions of coherence of ultra-high dimensional random matrices
- Variance Estimation Using Refitted Cross-validation in Ultrahigh Dimensional Regression
Cited by in corpus (7)
- Testing the martingale difference hypothesis in high dimension
- Spherical Cap Packing Asymptotics and Rank-Extreme Detection
- Cross-trait prediction accuracy of high-dimensional ridge-type estimators in genome-wide association studies
- Testing High Dimensional Mean Under Sparsity
- Sparse Sliced Inverse Regression Via Lasso
- Statistical Validity and Consistency of Big Data Analytics: A General Framework
- Cox reduction and confidence sets of models: a theoretical elucidation