Extensions of stability selection using subsamples of observations and covariates
arXiv:1407.4916 · doi:10.1007/s11222-015-9589-y
Abstract
We introduce extensions of stability selection, a method to stabilise variable selection methods introduced by Meinshausen and Bühlmann (J R Stat Soc 72:417-473, 2010). We propose to apply a base selection method repeatedly to random observation subsamples and covariate subsets under scrutiny, and to select covariates based on their selection frequency. We analyse the effects and benefits of these extensions. Our analysis generalizes the theoretical results of Meinshausen and Bühlmann (J R Stat Soc 72:417-473, 2010) from the case of half-samples to subsamples of arbitrary size. We study, in a theoretical manner, the effect of taking random covariate subsets using a simplified score model. Finally we validate these extensions on numerical experiments on both synthetic and real datasets, and compare the obtained results in detail to the original stability selection method.
accepted for publication in Statistics and Computing
References in corpus (5)
- Improving neural networks by preventing co-adaptation of feature detectors
- Lasso-type recovery of sparse representations for high-dimensional data
- Random lasso
- Correlated variables in regression: clustering and sparse estimation
- Sup-norm convergence rate and sign concentration property of Lasso and Dantzig estimators
Cited by in corpus (6)
- High-dimensional variable selection via low-dimensional adaptive learning
- A Novel Approach for Stable Selection of Informative Redundant Features from High Dimensional fMRI Data
- Heritability estimation in high dimensional linear mixed models
- Feature Selection for Huge Data via Minipatch Learning
- Bayesian Stability Selection and Inference on Selection Probabilities
- Pruning variable selection ensembles