Covariate Shift in High-Dimensional Random Feature Regression
arXiv:2111.08234
Abstract
A significant obstacle in the development of robust machine learning models is covariate shift, a form of distribution shift that occurs when the input distributions of the training and test sets differ while the conditional label distributions remain the same. Despite the prevalence of covariate shift in real-world applications, a theoretical understanding in the context of modern machine learning has remained lacking. In this work, we examine the exact high-dimensional asymptotics of random feature regression under covariate shift and present a precise characterization of the limiting test error, bias, and variance in this setting. Our results motivate a natural partial order over covariate shifts that provides a sufficient condition for determining when the shift will harm (or even help) test performance. We find that overparameterized models exhibit enhanced robustness to covariate shift, providing one of the first theoretical explanations for this intriguing phenomenon. Additionally, our analysis reveals an exact linear relationship between in-distribution and out-of-distribution generalization performance, offering an explanation for this surprising recent empirical observation.
107 pages, 10 figures
References in corpus (10)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Underspecification Presents Challenges for Credibility in Modern Machine Learning
- Do ImageNet Classifiers Generalize to ImageNet?
- The Neural Tangent Kernel in High Dimensions: Triple Descent and a Multi-Scale Theory of Generalization
- Accuracy on the Line: On the Strong Correlation Between Out-of-Distribution and In-Distribution Generalization
- Understanding Double Descent Requires a Fine-Grained Bias-Variance Decomposition
- What causes the test error? Going beyond bias-variance via ANOVA
- Provable Benefits of Overparameterization in Model Compression: From Double Descent to Pruning Neural Networks
- Near-Optimal Linear Regression under Distribution Shift
- Why do classifier accuracies show linear trends under distribution shift?