Random design analysis of ridge regression
arXiv:1106.2363
Abstract
This work gives a simultaneous analysis of both the ordinary least squares estimator and the ridge regression estimator in the random design setting under mild assumptions on the covariate/response distributions. In particular, the analysis provides sharp results on the ``out-of-sample'' prediction error, as opposed to the ``in-sample'' (fixed design) error. The analysis also reveals the effect of errors in the estimated covariance structure, as well as the effect of modeling errors, neither of which effects are present in the fixed design setting. The proofs of the main results are based on a simple decomposition lemma combined with concentration inequalities for random vectors and matrices.
Cited by in corpus (41)
- Divide and Conquer Kernel Ridge Regression: A Distributed Algorithm with Minimax Optimal Rates
- Exploiting Shared Representations for Personalized Federated Learning
- The lower tail of random quadratic forms, with applications to ordinary least squares and restricted eigenvalue properties
- Few-Shot Learning via Learning the Representation, Provably
- Provable Meta-Learning of Linear Representations
- Predicting What You Already Know Helps: Provable Self-Supervised Learning
- What are the Statistical Limits of Offline RL with Linear Function Approximation?
- Correlated random features for fast semi-supervised learning
- Subspace Embeddings and -Regression Using Exponential Random Variables
- Active Regression by Stratification
- A Continuous-Time View of Early Stopping for Least Squares
- Learning with incremental iterative regularization
- Minimax Linear Estimation of the Retargeted Mean
- Generalization Error Rates in Kernel Regression: The Crossover from the Noiseless to Noisy Regime
- In-N-Out: Pre-Training and Self-Training using Auxiliary Information for Out-of-Distribution Robustness
- Kernel Ridge Regression via Partitioning
- Affine Invariant Covariance Estimation for Heavy-Tailed Distributions
- Kernel Truncated Randomized Ridge Regression: Optimal Rates and Low Noise Acceleration
- Hierarchically Regularized Deep Forecasting
- Learning Some Popular Gaussian Graphical Models without Condition Number Bounds
- Error Scaling Laws for Kernel Classification under Source and Capacity Conditions
- The Benefits of Implicit Regularization from SGD in Least Squares Problems
- Stochastic Online Optimization using Kalman Recursion
- Analysis of a randomized approximation scheme for matrix multiplication
- Adversarially Robust Estimate and Risk Analysis in Linear Regression
- Optimal Rates for Learning with Nyström Stochastic Gradient Methods
- Taming heavy-tailed features by shrinkage
- Heteroscedasticity-aware residuals-based contextual stochastic optimization
- Meta-learning for mixed linear regression
- Single Point Transductive Prediction
- On the benefits of maximum likelihood estimation for Regression and Forecasting
- Revisiting minimum description length complexity in overparameterized models
- Naive imputation implicitly regularizes high-dimensional linear models
- Geometric Exploration for Online Control
- Regularized Loss Minimizers with Local Data Perturbation: Consistency and Data Irrecoverability
- `Basic' Generalization Error Bounds for Least Squares Regression with Well-specified Models
- ParK: Sound and Efficient Kernel Ridge Regression by Feature Space Partitions
- Conditionally Strongly Log-Concave Generative Models
- Curvature-Exploiting Acceleration of Elastic Net Computations
- Minimax Optimal Regression over Sobolev Spaces via Laplacian Eigenmaps on Neighborhood Graphs
- Sparsity in Partially Controllable Linear Systems