Robust linear least squares regression
arXiv:1010.0074 · doi:10.1214/11-AOS918
Abstract
We consider the problem of robustly predicting as well as the best linear combination of given functions in least squares regression, and variants of this problem including constraints on the parameters of the linear combination. For the ridge estimator and the ordinary least squares estimator, and their variants, we provide new risk bounds of order without logarithmic factor unlike some standard results, where is the size of the training data. We also provide a new estimator with better deviations in the presence of heavy-tailed noise. It is based on truncating differences of losses in a min--max framework and satisfies a risk bound both in expectation and in deviations. The key common surprising factor of these results is the absence of exponential moment condition on the output distribution while achieving exponential deviations. All risk bounds are obtained through a PAC-Bayesian analysis on truncated differences of losses. Experimental results strongly back up our truncated min--max estimator.
Published in at http://dx.doi.org/10.1214/11-AOS918 the Annals of Statistics (http://www.imstat.org/aos/) by the Institute of Mathematical Statistics (http://www.imstat.org). arXiv admin note: significant text overlap with arXiv:0902.1733
References in corpus (1)
Cited by in corpus (13)
- Geometric median and robust estimation in Banach spaces
- Robust linear least squares regression
- Empirical risk minimization for heavy-tailed losses
- Simpler PAC-Bayesian Bounds for Hostile Data
- A new method for estimation and model selection: -estimation
- Exact minimax risk for linear least squares, and the lower tail of sample covariance matrices
- Bayesian methods for low-rank matrix estimation: short survey and theoretical study
- PAC-Bayesian Estimation and Prediction in Sparse Additive Models
- Benchmarking Particle Filter Algorithms for Efficient Velodyne-Based Vehicle Localization
- Performance of empirical risk minimization in linear aggregation
- Empirical risk minimization is optimal for the convex aggregation problem
- Performance of Empirical Risk Minimization for Linear Regression with Dependent Data
- Trimmed sample means for robust uniform mean estimation and regression