Optimal ridge penalty for real-world high-dimensional data can be zero or negative due to the implicit ridge regularization
arXiv:1805.10939
Abstract
A conventional wisdom in statistical learning is that large models require strong regularization to prevent overfitting. Here we show that this rule can be violated by linear regression in the underdetermined situation under realistic conditions. Using simulations and real-life high-dimensional data sets, we demonstrate that an explicit positive ridge penalty can fail to provide any improvement over the minimum-norm least squares estimator. Moreover, the optimal value of ridge penalty in this situation can be negative. This happens when the high-variance directions in the predictor space can predict the response variable, which is often the case in the real-world high-dimensional data. In this regime, low-variance directions provide an implicit ridge regularization and can make any further positive ridge penalty detrimental. We prove that augmenting any linear model with random covariates and using minimum-norm estimator is asymptotically equivalent to adding the ridge penalty. We use a spiked covariance model as an analytically tractable example and prove that the optimal ridge penalty in this case is negative when .
References in corpus (8)
- Reconciling modern machine learning practice and the bias-variance trade-off
- Benign Overfitting in Linear Regression
- The generalization error of random features regression: Precise asymptotics and double descent curve
- Overfitting or perfect fitting? Risk bounds for classification and regression rules that interpolate
- Implicit Regularization in Deep Learning
- Theory of Deep Learning III: explaining the non-overfitting puzzle
- More Data Can Hurt for Linear Regression: Sample-wise Double Descent
- A New Look at an Old Problem: A Universal Learning Approach to Linear Regression
Cited by in corpus (9)
- The generalization error of max-margin linear classifiers: Benign overfitting and high dimensional asymptotics in the overparametrized regime
- Benign overfitting in ridge regression
- Optimal Regularization Can Mitigate Double Descent
- Exact expressions for double descent and implicit regularization via surrogate random design
- On the Inherent Regularization Effects of Noise Injection During Training
- On the interplay between data structure and loss function in classification problems
- Improving benchmarks for autonomous vehicles testing using synthetically generated images
- Adaptive Reference-Guided Estimation of Principal Component Subspace in High Dimensions
- Ridge-penalized adaptive Mantel test and its application in imaging genetics