A scalable estimate of the extra-sample prediction error via approximate leave-one-out
arXiv:1801.10243
Abstract
The paper considers the problem of out-of-sample risk estimation under the high dimensional settings where standard techniques such as -fold cross validation suffer from large biases. Motivated by the low bias of the leave-one-out cross validation (LO) method, we propose a computationally efficient closed-form approximate leave-one-out formula (ALO) for a large class of regularized estimators. Given the regularized estimate, calculating ALO requires minor computational overhead. With minor assumptions about the data generating process, we obtain a finite-sample upper bound for . Our theoretical analysis illustrates that with overwhelming probability, when , where the dimension of the feature vectors may be comparable with or even greater than the number of observations, . Despite the high-dimensionality of the problem, our theoretical results do not require any sparsity assumption on the vector of regression coefficients. Our extensive numerical experiments show that decreases as increase, revealing the excellent finite sample performance of ALO. We further illustrate the usefulness of our proposed out-of-sample risk estimation method by an example of real recordings from spatially sensitive neurons (grid cells) in the medial entorhinal cortex of a rat.
References in corpus (4)
- On the "degrees of freedom" of the lasso
- Correlations and functional connections in a population of grid cells
- Evaluation and selection of models for out-of-sample prediction when the sample size is small relative to the complexity of the data-generating process
- Conditional predictive inference post model selection
Cited by in corpus (13)
- On the Accuracy of Influence Functions for Measuring Group Effects
- Approximate Leave-One-Out for Fast Parameter Tuning in High Dimensions
- A Higher-Order Swiss Army Infinitesimal Jackknife
- Approximate Leave-One-Out for High-Dimensional Non-Differentiable Learning Problems
- Leave Zero Out: Towards a No-Cross-Validation Approach for Model Selection
- Error bounds in estimating the out-of-sample prediction error using leave-one-out cross validation in high-dimensions
- Consistent Risk Estimation in Moderately High-Dimensional Linear Regression
- Approximate Cross-Validation for Structured Models
- Can we globally optimize cross-validation loss? Quasiconvexity in ridge regression
- Optimizing Approximate Leave-one-out Cross-validation to Tune Hyperparameters
- Accelerating Cross-Validation in Multinomial Logistic Regression with -Regularization
- Approximate Cross-Validation with Low-Rank Data in High Dimensions
- Approximate Cross-validation: Guarantees for Model Assessment and Selection