Consistent Risk Estimation in Moderately High-Dimensional Linear Regression
arXiv:1902.01753
Abstract
Risk estimation is at the core of many learning systems. The importance of this problem has motivated researchers to propose different schemes, such as cross validation, generalized cross validation, and Bootstrap. The theoretical properties of such estimates have been extensively studied in the low-dimensional settings, where the number of predictors is much smaller than the number of observations . However, a unifying methodology accompanied with a rigorous theory is lacking in high-dimensional settings. This paper studies the problem of risk estimation under the moderately high-dimensional asymptotic setting and ( is a fixed number), and proves the consistency of three risk estimates that have been successful in numerical studies, i.e., leave-one-out cross validation (LOOCV), approximate leave-one-out (ALO), and approximate message passing (AMP)-based techniques. A corner stone of our analysis is a bound that we obtain on the discrepancy of the `residuals' obtained from AMP and LOOCV. This connection not only enables us to obtain a more refined information on the estimates of AMP, ALO, and LOOCV, but also offers an upper bound on the convergence rate of each estimate.
References in corpus (12)
- Message Passing Algorithms for Compressed Sensing
- Understanding Black-box Predictions via Influence Functions
- Templates for Convex Cone Problems with Applications to Sparse Signal Recovery
- Benign Overfitting in Linear Regression
- Consistency of cross validation for comparing regression procedures
- Geometric Inference for General High-Dimensional Linear Inverse Problems
- Asymptotic Analysis of LASSOs Solution Path with Implications for Approximate Message Passing
- A scalable estimate of the extra-sample prediction error via approximate leave-one-out
- Variance Breakdown of Huber (M)-estimators:
- Robustness in sparse linear models: relative efficiency based on robust approximate message passing
- Approximate Leave-One-Out for High-Dimensional Non-Differentiable Learning Problems
- Optimization-based AMP for Phase Retrieval: The Impact of Initialization and -regularization
Cited by in corpus (4)
- Error bounds in estimating the out-of-sample prediction error using leave-one-out cross validation in high-dimensions
- Out-of-sample error estimate for robust M-estimators with convex penalty
- Convex and Nonconvex Optimization Are Both Minimax-Optimal for Noisy Blind Deconvolution under Random Designs
- Comparing Classes of Estimators: When does Gradient Descent Beat Ridge Regression in Linear Models?