The Lasso with general Gaussian designs with applications to hypothesis testing
arXiv:2007.13716
Abstract
The Lasso is a method for high-dimensional regression, which is now commonly used when the number of covariates is of the same order or larger than the number of observations . Classical asymptotic normality theory does not apply to this model due to two fundamental reasons: The regularized risk is non-smooth; The distance between the estimator and the true parameters vector cannot be neglected. As a consequence, standard perturbative arguments that are the traditional basis for asymptotic normality fail. On the other hand, the Lasso estimator can be precisely characterized in the regime in which both and are large and is of order one. This characterization was first obtained in the case of Gaussian designs with i.i.d. covariates: here we generalize it to Gaussian correlated designs with non-singular covariance structure. This is expressed in terms of a simpler ``fixed-design'' model. We establish non-asymptotic bounds on the distance between the distribution of various quantities in the two models, which hold uniformly over signals in a suitable sparsity class and over values of the regularization parameter. As an application, we study the distribution of the debiased Lasso and show that a degrees-of-freedom correction is necessary for computing valid confidence intervals.
final version accepted to Annals of Statistics
References in corpus (3)
Cited by in corpus (17)
- Learning curves of generic features maps for realistic datasets with a teacher-student model
- Gaussian Universality of Perceptrons with Random Labels
- A Power Analysis of the Conditional Randomization Test and Knockoffs
- Inference for Heteroskedastic PCA with Missing Data
- Derivatives and residual distribution of regularized M-estimators with application to adaptive tuning
- Out-of-sample error estimate for robust M-estimators with convex penalty
- CAD: Debiasing the Lasso with inaccurate covariate model
- Local convexity of the TAP free energy and AMP convergence for Z2-synchronization
- Randomized tests for high-dimensional regression: A more efficient and powerful solution
- Minimum -norm interpolators: Precise asymptotics and multiple descent
- Learning Gaussian Mixtures with Generalised Linear Models: Precise Asymptotics in High-dimensions
- Chi-square and normal inference in high-dimensional multi-task regression
- Efficient Designs of SLOPE Penalty Sequences in Finite Dimension
- A New Perspective on Debiasing Linear Regressions
- Asymptotic normality of robust -estimators with convex penalty
- Characterizing the SLOPE Trade-off: A Variational Perspective and the Donoho-Tanner Limit
- Sharp global convergence guarantees for iterative nonconvex optimization: A Gaussian process perspective