Cross-validation: what does it estimate and how well does it do it?
arXiv:2104.00673 · doi:10.1080/01621459.2023.2197686
Abstract
Cross-validation is a widely-used technique to estimate prediction error, but its behavior is complex and not fully understood. Ideally, one would like to think that cross-validation estimates the prediction error for the model at hand, fit to the training data. We prove that this is not the case for the linear model fit by ordinary least squares; rather it estimates the average prediction error of models fit on other unseen training sets drawn from the same population. We further show that this phenomenon occurs for most popular estimates of prediction error, including data splitting, bootstrapping, and Mallow's Cp. Next, the standard confidence intervals for prediction error derived from cross-validation may have coverage far below the desired level. Because each data point is used for both training and testing, there are correlations among the measured accuracies for each fold, and so the usual estimate of variance is too small. We introduce a nested cross-validation scheme to estimate this variance more accurately, and we show empirically that this modification leads to intervals with approximately correct coverage in many examples where traditional cross-validation intervals fail.
References in corpus (3)
Cited by in corpus (15)
- Relating the Partial Dependence Plot and Permutation Feature Importance to the Data Generating Process
- On Leakage in Machine Learning Pipelines
- Empirical investigation of multi-source cross-validation in clinical ECG classification
- The out-of-sample : estimation and inference
- Integration of nested cross-validation, automated hyperparameter optimization, high-performance computing to reduce and quantify the variance of test performance estimation of deep learning models
- Deep reinforcement learning for smart calibration of radio telescopes
- A New Formula for Faster Computation of the K-Fold Cross-Validation and Good Regularisation Parameter Values in Ridge Regression
- Uncertainty in Bayesian Leave-One-Out Cross-Validation Based Model Comparison
- Model-free Bootstrap and Conformal Prediction in Regression: Conditionality, Conjecture Testing, and Pertinent Prediction Intervals
- Concentration Inequalities for Cross-validation in Scattered Data Approximation
- Weighted Leave-One-Out Cross Validation
- Stable and Robust Hyper-Parameter Selection Via Robust Information Sharing Cross-Validation
- The Delaunay Density Diagnostic
- Predicting Loss Risks for B2B Tendering Processes
- Large multi-response linear regression estimation based on low-rank pre-smoothing