Variable selection with error control: Another look at Stability Selection
arXiv:1105.5578 · doi:10.1111/j.1467-9868.2011.01034.x
Abstract
Stability Selection was recently introduced by Meinshausen and Buhlmann (2010) as a very general technique designed to improve the performance of a variable selection algorithm. It is based on aggregating the results of applying a selection procedure to subsamples of the data. We introduce a variant, called Complementary Pairs Stability Selection (CPSS), and derive bounds both on the expected number of variables included by CPSS that have low selection probability under the original procedure, and on the expected number of high selection probability variables that are excluded. These results require no (e.g. exchangeability) assumptions on the underlying model or on the quality of the original selection procedure. Under reasonable shape restrictions, the bounds can be further tightened, yielding improved error control, and therefore increasing the applicability of the methodology.
25 pages, 9 figures
References in corpus (2)
Cited by in corpus (28)
- On asymptotically optimal confidence regions and tests for high-dimensional models
- High-Dimensional Inference: Confidence Intervals, -Values and R-Software hdi
- Learning hydrodynamic equations for active matter from particle simulations and experiments
- Stability selection for component-wise gradient boosting in multiple dimensions
- Extending Statistical Boosting - An Overview of Recent Methodological Developments
- Network classification with applications to brain connectomics
- A Unified Framework of Constrained Regression
- Goodness of fit tests for high-dimensional linear models
- Automated calibration for stability selection in penalised regression and graphical models
- Graphical Models for Zero-Inflated Single Cell Gene Expression
- Extensions of stability selection using subsamples of observations and covariates
- Hierarchical inference for genome-wide association studies: a view on methodology with software
- High-dimensional variable selection via low-dimensional adaptive learning
- Randomized maximum-contrast selection: subagging for large-scale regression
- Rank-transformed subsampling: inference for multiple data splitting and exchangeable p-values
- View selection in multi-view stacking: Choosing the meta-learner
- Network-Guided Biomarker Discovery
- Infinite-Dimensional Sparse Learning in Linear System Identification
- Employing an Adjusted Stability Measure for Multi-Criteria Model Fitting on Data Sets with Similar Features
- Nonparametric causal structure learning in high dimensions
- Bayesian Stability Selection and Inference on Selection Probabilities
- FRI -- Feature Relevance Intervals for Interpretable and Interactive Data Exploration
- Flexible variable selection in the presence of missing data
- Integrated path stability selection
- Replica Analysis for Ensemble Techniques in Variable Selection
- Nonparametric IPSS: Fast, flexible feature selection with false discovery control
- Fast variable selection for distributional regression with application to continuous glucose monitoring data
- Cross-Channel Unlabeled Sensing over a Union of Signal Subspaces