Exponential Screening and optimal rates of sparse estimation
arXiv:1003.2654
Abstract
In high-dimensional linear regression, the goal pursued here is to estimate an unknown regression function using linear combinations of a suitable set of covariates. One of the key assumptions for the success of any statistical procedure in this setup is to assume that the linear combination is sparse in some sense, for example, that it involves only few covariates. We consider a general, non necessarily linear, regression with Gaussian noise and study a related question that is to find a linear combination of approximating functions, which is at the same time sparse and has small mean squared error (MSE). We introduce a new estimation procedure, called Exponential Screening that shows remarkable adaptation properties. It adapts to the linear combination that optimally balances MSE and sparsity, whether the latter is measured in terms of the number of non-zero entries in the combination ( norm) or in terms of the global weight of the combination ( norm). The power of this adaptation result is illustrated by showing that Exponential Screening solves optimally and simultaneously all the problems of aggregation in Gaussian regression that have been discussed in the literature. Moreover, we show that the performance of the Exponential Screening estimator cannot be improved in a minimax sense, even if the optimal sparsity is known in advance. The theoretical and numerical superiority of Exponential Screening compared to state-of-the-art sparse procedures is also discussed.
References in corpus (3)
Cited by in corpus (42)
- Near-linear time approximation algorithms for optimal transport via Sinkhorn iteration
- Approximation and Estimation for High-Dimensional Deep Learning Networks
- Prediction error of cross-validated Lasso
- Sparse Nonlinear Regression: Parameter Estimation and Asymptotic Inference
- A General Framework for Bayes Structured Linear Models
- On Robustness of Principal Component Regression
- Trust, but verify: benefits and pitfalls of least-squares refitting in high dimensions
- Transfer Learning for High-dimensional Linear Regression: Prediction, Estimation, and Minimax Optimality
- Regularization and the small-ball method II: complexity dependent error rates
- Sparsity regret bounds for individual sequences in online linear regression
- An explicit analysis of the entropic penalty in linear programming
- An l1-Oracle Inequality for the Lasso
- Adaptive Minimax Estimation over Sparse -Hulls
- Volume Ratio, Sparsity, and Minimaxity under Unitarily Invariant Norms
- Estimation of Covariance Matrices under Sparsity Constraints
- Optimal Estimation and Prediction for Dense Signals in High-Dimensional Linear Models
- General framework for projection structures
- Variable Selection with Exponential Weights and -Penalization
- Adaptive Estimation in Structured Factor Models with Applications to Overlapping Clustering
- Group Regularized Estimation under Structural Hierarchy
- Sparse Partially Linear Additive Models
- Alternating minimization for dictionary learning: Local Convergence Guarantees
- High-dimensional Log-Error-in-Variable Regression with Applications to Microbial Compositional Data Analysis
- Aggregation of supports along the Lasso path
- Robust Reduced Rank Regression
- Aggregation of Affine Estimators
- Note on Existence and Non-Existence of Large Subsets of Binary Vectors with Similar Distances
- High-dimensional Adaptive Minimax Sparse Estimation with Interactions
- High-dimensional instrumental variables regression and confidence sets -- v2/2012
- Maximum Regularized Likelihood Estimators: A General Prediction Theory and Applications
- Inference Without Compatibility
- Block based refitting in sparse regularisation
- On Cross-validation for Sparse Reduced Rank Regression
- Statistical Inference for Data-adaptive Doubly Robust Estimators with Survival Outcomes
- Refitting solutions promoted by sparse analysis regularization with block penalties
- An \ell_1-oracle inequality for the Lasso in finite mixture of multivariate Gaussian regression models
- Graphical Exponential Screening
- Robust Orthogonal Complement Principal Component Analysis
- Sparse additive regression on a regular lattice
- On prediction with the LASSO when the design is not incoherent
- Estimation of a sparse group of sparse vectors
- Estimating the Random Effect in Big Data Mixed Models