Least squares after model selection in high-dimensional sparse models
arXiv:1001.0188 · doi:10.3150/11-BEJ410
Abstract
In this article we study post-model selection estimators that apply ordinary least squares (OLS) to the model selected by first-step penalized estimators, typically Lasso. It is well known that Lasso can estimate the nonparametric regression function at nearly the oracle rate, and is thus hard to improve upon. We show that the OLS post-Lasso estimator performs at least as well as Lasso in terms of the rate of convergence, and has the advantage of a smaller bias. Remarkably, this performance occurs even if the Lasso-based model selection "fails" in the sense of missing some components of the "true" regression model. By the "true" model, we mean the best s-dimensional approximation to the nonparametric regression function chosen by the oracle. Furthermore, OLS post-Lasso estimator can perform strictly better than Lasso, in the sense of a strictly faster rate of convergence, if the Lasso-based model selection correctly includes all components of the "true" model as a subset and also achieves sufficient sparsity. In the extreme case, when Lasso perfectly selects the "true" model, the OLS post-Lasso estimator becomes the oracle estimator. An important ingredient in our analysis is a new sparsity bound on the dimension of the model selected by Lasso, which guarantees that this dimension is at most of the same order as the dimension of the "true" model. Our rate results are nonasymptotic and hold in both parametric and nonparametric models. Moreover, our analysis is not limited to the Lasso estimator acting as a selector in the first step, but also applies to any other estimator, for example, various forms of thresholded Lasso, with good rates and good sparsity properties. Our analysis covers both traditional thresholding and a new practical, data-driven thresholding scheme that induces additional sparsity subject to maintaining a certain goodness of fit. The latter scheme has theoretical guarantees similar to those of Lasso or OLS post-Lasso, but it dominates those procedures as well as traditional thresholding in a wide variety of experiments.
Published in at http://dx.doi.org/10.3150/11-BEJ410 the Bernoulli (http://isi.cbs.nl/bernoulli/) by the International Statistical Institute/Bernoulli Society (http://isi.cbs.nl/BS/bshome.htm)
References in corpus (9)
- Asymptotic properties of bridge estimators in sparse high-dimensional regression models
- The sparsity and bias of the Lasso selection in high-dimensional linear regression
- Lasso-type recovery of sparse representations for high-dimensional data
- High-dimensional generalized linear models and the lasso
- Sparsity oracle inequalities for the Lasso
- Aggregation for Gaussian regression
- Sharp thresholds for high-dimensional and noisy recovery of sparsity
- Sup-norm convergence rate and sign concentration property of Lasso and Dantzig estimators
- Consistent selection via the Lasso for high dimensional approximating regression models
Cited by in corpus (73)
- On asymptotically optimal confidence regions and tests for high-dimensional models
- Confidence Intervals and Hypothesis Testing for High-Dimensional Regression
- Robust Inference on Average Treatment Effects with Possibly More Covariates than Observations
- Gaussian approximations and multiplier bootstrap for maxima of sums of high-dimensional random vectors
- Uniform Post Selection Inference for LAD Regression and Other Z-estimation problems
- Lasso adjustments of treatment effect estimates in randomized experiments
- Valid Post-Selection and Post-Regularization Inference: An Elementary, General Approach
- Pivotal estimation via square-root Lasso in nonparametric regression
- High-dimensional instrumental variables regression and confidence sets
- Endogeneity in high dimensions
- Double Machine Learning based Program Evaluation under Unconfoundedness
- Heterogeneous Employment Effects of Job Search Programmes: A Machine Learning Approach
- Moving Beyond Sub-Gaussianity in High-Dimensional Statistics: Applications in Covariance Estimation and Linear Regression
- Double/Debiased Machine Learning for Treatment and Causal Parameters
- Does Data Splitting Improve Prediction?
- Distributed High-dimensional Regression Under a Quantile Loss Function
- When are Google data useful to nowcast GDP? An approach via pre-selection and shrinkage
- Nonparametric estimation of causal heterogeneity under high-dimensional confounding
- BigVAR: Tools for Modeling Sparse High-Dimensional Multivariate Time Series
- Efficient Smoothed Concomitant Lasso Estimation for High Dimensional Regression
- Trust, but verify: benefits and pitfalls of least-squares refitting in high dimensions
- Statistical Inferences for Polarity Identification in Natural Language
- De-Biasing The Lasso With Degrees-of-Freedom Adjustment
- A review of regularised estimation methods and cross-validation in spatiotemporal statistics
- High-dimensional empirical likelihood inference
- Sparse structures with LASSO through Principal Components: forecasting GDP components in the short-run
- High-Dimensional Sparse Linear Bandits
- Optimal bounds for aggregation of affine estimators
- Uniformly Valid Post-Regularization Confidence Regions for Many Functional Parameters in Z-Estimation Framework
- Random weighting in LASSO regression
- On estimation of the diagonal elements of a sparse precision matrix
- High-dimensional variable selection via low-dimensional adaptive learning
- On the Generalization Properties of Adversarial Training
- Sparse Feature Selection Makes Batch Reinforcement Learning More Sample Efficient
- Propensity Score Weighting for Causal Subgroup Analysis
- Double Machine Learning for Partially Linear Mixed-Effects Models with Repeated Measurements
- Inference for high-dimensional instrumental variables regression
- Two Stage Non-penalized Corrected Least Squares for High Dimensional Linear Models with Measurement error or Missing Covariates
- A Projection Based Conditional Dependence Measure with Applications to High-dimensional Undirected Graphical Models
- Two-step estimation of high dimensional additive models
- Efficient Difference-in-Differences Estimation with High-Dimensional Common Trend Confounding
- Asymptotic Properties of Lasso+mLS and Lasso+Ridge in Sparse High-dimensional Linear Regression
- Two-sample inference for high-dimensional Markov networks
- Omitted variable bias of Lasso-based inference methods: A finite sample analysis
- Finite mixture regression: A sparse variable selection by model selection for clustering
- Regularizing Double Machine Learning in Partially Linear Endogenous Models
- Non-Bayesian Post-Model-Selection Estimation as Estimation Under Model Misspecification
- Ultra high dimensional generalized additive model: Unified Theory and Methods
- Treatment Effect Estimation with Observational Network Data using Machine Learning
- Panel Data Quantile Regression with Grouped Fixed Effects
- A Lasso-OLS Hybrid Approach to Covariate Selection and Average Treatment Effect Estimation for Clustered RCTs Using Design-Based Methods
- Sharp Inference on Selected Subgroups in Observational Studies
- Predicting Exporters with Machine Learning
- Learning low dimensional word based linear classifiers using Data Shared Adaptive Bootstrap Aggregated Lasso with application to IMDb data
- Post Selection Shrinkage Estimation for High Dimensional Data Analysis
- Macroeconomic Forecasting and Variable Selection with a Very Large Number of Predictors: A Penalized Regression Approach
- Data augmentation for non-Gaussian regression models using variance-mean mixtures
- Post-selection inference on high-dimensional varying-coefficient quantile regression model
- Personalised dynamic super learning: an application in predicting hemodiafiltration convection volumes
- Two step estimations via the Dantzig selector for models of stochastic processes with high-dimensional parameters
- Dynamic mode decomposition for detecting oscillatory transient activity via sparsity and smoothness regularization
- Model Determination for High-Dimensional Longitudinal Data with Missing Observations: An Application to Microfinance Data
- DebiNet: Debiasing Linear Models with Nonlinear Overparameterized Neural Networks
- Block based refitting in sparse regularisation
- Sparse Regularization in Marketing and Economics
- Nested Model Averaging on Solution Path for High-dimensional Linear Regression
- Bayesian causal inference with some invalid instrumental variables
- Adaptive Discrete Smoothing for High-Dimensional and Nonlinear Panel Data
- HCmodelSets: An R package for specifying sets of well-fitting models in regression with a large number of potential explanatory variables
- Multiple imputation and selection of ordinal level 2 predictors in multilevel models. An analysis of the relationship between student ratings and teacher beliefs and practices
- Sparse regression with highly correlated predictors
- Uniform-in-Submodel Bounds for Linear Regression in a Model Free Framework
- Selecting Penalty Parameters of High-Dimensional M-Estimators using Bootstrapping after Cross-Validation