The sparsity and bias of the Lasso selection in high-dimensional linear regression
arXiv:0808.0967 · doi:10.1214/07-AOS520
Abstract
Meinshausen and Buhlmann [Ann. Statist. 34 (2006) 1436--1462] showed that, for neighborhood selection in Gaussian graphical models, under a neighborhood stability condition, the LASSO is consistent, even when the number of variables is of greater order than the sample size. Zhao and Yu [(2006) J. Machine Learning Research 7 2541--2567] formalized the neighborhood stability condition in the context of linear regression as a strong irrepresentable condition. That paper showed that under this condition, the LASSO selects exactly the set of nonzero regression coefficients, provided that these coefficients are bounded away from zero at a certain rate. In this paper, the regression coefficients outside an ideal model are assumed to be small, but not necessarily zero. Under a sparse Riesz condition on the correlation of design variables, we prove that the LASSO selects a model of the correct order of dimensionality, controls the bias of the selected model at a level determined by the contributions of small regression coefficients and threshold bias, and selects all coefficients of greater order than the bias of the selected model. Moreover, as a consequence of this rate consistency of the LASSO in model selection, it is proved that the sum of error squares for the mean response and the -loss for the regression coefficients converge at the best possible rates under the given conditions. An interesting aspect of our results is that the logarithm of the number of variables can be of the same order as the sample size for certain random dependent designs.
Published in at http://dx.doi.org/10.1214/07-AOS520 the Annals of Statistics (http://www.imstat.org/aos/) by the Institute of Mathematical Statistics (http://www.imstat.org)
References in corpus (5)
- Lasso-type recovery of sparse representations for high-dimensional data
- High-dimensional generalized linear models and the lasso
- Sparsity oracle inequalities for the Lasso
- Sharp thresholds for high-dimensional and noisy recovery of sparsity
- Discussion: The Dantzig selector: Statistical estimation when is much larger than
Cited by in corpus (28)
- Nearly unbiased variable selection under minimax concave penalty
- Variable selection in nonparametric additive models
- Lasso-type recovery of sparse representations for high-dimensional data
- Performance guarantees for individualized treatment rules
- The Emerging Trends of Multi-Label Learning
- High-dimensional variable selection
- Optimal rates of convergence for covariance matrix estimation
- A Selective Review of Group Selection in High-Dimensional Models
- Needles and Straw in a Haystack: Posterior concentration for possibly sparse sequences
- Two sample tests for high-dimensional covariance matrices
- Least angle and penalized regression: A review
- L1-Penalization for Mixture Regression Models
- SCAD-penalized regression in high-dimensional partially linear models
- Variable selection and regression analysis for graph-structured covariates with an application to genomics
- Estimation for High-Dimensional Linear Mixed-Effects Models Using -Penalization
- Some sharp performance bounds for least squares regression with regularization
- High-dimensional variable selection via tilting
- Consistent group selection in high-dimensional linear regression
- Discussion: One-step sparse estimates in nonconcave penalized likelihood models
- Estimation in high-dimensional linear models with deterministic design matrices
- Parametric or nonparametric? A parametricness index for model selection
- Discussion: One-step sparse estimates in nonconcave penalized likelihood models
- A general theory of regression adjustment for covariate-adaptive randomization: OLS, Lasso, and beyond
- Robust Information Criterion for Model Selection in Sparse High-Dimensional Linear Regression Models
- Culling the herd of moments with penalized empirical likelihood
- Smoothing the Edges: Smooth Optimization for Sparse Regularization using Hadamard Overparametrization
- KL-BSS: Rethinking optimality for neighbourhood selection in structural equation models
- Weighted residual empirical processes, martingale transformations, and model specification tests for regressions with diverging number of parameters