Impacts of high dimensionality in finite samples
arXiv:1311.2742 · doi:10.1214/13-AOS1149
Abstract
High-dimensional data sets are commonly collected in many contemporary applications arising in various fields of scientific research. We present two views of finite samples in high dimensions: a probabilistic one and a nonprobabilistic one. With the probabilistic view, we establish the concentration property and robust spark bound for large random design matrix generated from elliptical distributions, with the former related to the sure screening property and the latter related to sparse model identifiability. An interesting concentration phenomenon in high dimensions is revealed. With the nonprobabilistic view, we derive general bounds on dimensionality with some distance constraint on sparse models. These results provide new insights into the impacts of high dimensionality in finite samples.
Published in at http://dx.doi.org/10.1214/13-AOS1149 the Annals of Statistics (http://www.imstat.org/aos/) by the Institute of Mathematical Statistics (http://www.imstat.org)
References in corpus (4)
- On the conditions used to prove oracle results for the Lasso
- High-dimensional classification using features annealed independence rules
- Statistical Challenges with High Dimensionality: Feature Selection in Knowledge Discovery
- A unified approach to model selection and sparse recovery using regularized least squares