Variable selection in high-dimensional linear models: partially faithful distributions and the PC-simple algorithm
arXiv:0906.3204 · doi:10.1093/biomet/asq008
Abstract
We consider variable selection in high-dimensional linear models where the number of covariates greatly exceeds the sample size. We introduce the new concept of partial faithfulness and use it to infer associations between the covariates and the response. Under partial faithfulness, we develop a simplified version of the PC algorithm (Spirtes et al., 2000), the PC-simple algorithm, which is computationally feasible even with thousands of covariates and provides consistent variable selection under conditions on the random design matrix that are of a different nature than coherence conditions for penalty-based approaches like the Lasso. Simulations and application to real data show that our method is competitive compared to penalty-based approaches. We provide an efficient implementation of the algorithm in the R-package pcalg.
20 pages, 3 figures
References in corpus (6)
- Simultaneous analysis of Lasso and Dantzig selector
- The sparsity and bias of the Lasso selection in high-dimensional linear regression
- Lasso-type recovery of sparse representations for high-dimensional data
- High-dimensional generalized linear models and the lasso
- High-dimensional variable selection
- Stability Selection
Cited by in corpus (21)
- High-dimensional additive modeling
- Quantile-adaptive model-free variable screening for high-dimensional heterogeneous data
- Statistical significance in high-dimensional linear models
- Endogeneity in high dimensions
- High-dimensional variable selection via tilting
- Causal Decision Trees
- Revisiting Marginal Regression
- High-dimensional variable selection via low-dimensional adaptive learning
- Evaluating structure learning algorithms with a balanced scoring function
- Two Stage Non-penalized Corrected Least Squares for High Dimensional Linear Models with Measurement error or Missing Covariates
- Causal query in observational data with hidden variables
- Stable Prediction via Leveraging Seed Variable
- Bayesian Ultrahigh-Dimensional Screening Via MCMC
- Markov Neighborhood Regression for High-Dimensional Inference
- Consistency of Bayesian Linear Model Selection With a Growing Number of Parameters
- Identify treatment effect patterns for personalised decisions
- Variable Selection for Survival Data with A Class of Adaptive Elastic Net Techniques
- Generalized Fiducial Inference for Ultrahigh Dimensional Regression
- KL-BSS: Rethinking optimality for neighbourhood selection in structural equation models
- Cox reduction and confidence sets of models: a theoretical elucidation
- High-dimensional Statistical Inference and Variable Selection Using Sufficient Dimension Association