On Bayes Risk Lower Bounds
arXiv:1410.0503
Abstract
This paper provides a general technique for lower bounding the Bayes risk of statistical estimation, applicable to arbitrary loss functions and arbitrary prior distributions. A lower bound on the Bayes risk not only serves as a lower bound on the minimax risk, but also characterizes the fundamental limit of any estimator given the prior knowledge. Our bounds are based on the notion of -informativity, which is a function of the underlying class of probability measures and the prior. Application of our bounds requires upper bounds on the -informativity, thus we derive new upper bounds on -informativity which often lead to tight Bayes risk lower bounds. Our technique leads to generalizations of a variety of classical minimax bounds (e.g., generalized Fano's inequality). Our Bayes risk lower bounds can be directly applied to several concrete estimation problems, including Gaussian location models, generalized linear models, and principal component analysis for spiked covariance models. To further demonstrate the applications of our Bayes risk lower bounds to machine learning problems, we present two new theoretical results: (1) a precise characterization of the minimax risk of learning spherical Gaussian mixture models under the smoothed analysis framework, and (2) lower bounds for the Bayes risk under a natural prior for both the prediction and estimation errors for high-dimensional sparse linear regression under an improper learning setting.
55 pages, 2 figures
References in corpus (4)
Cited by in corpus (8)
- -divergence Inequalities
- Provable Meta-Learning of Linear Representations
- Reward-Free Exploration for Reinforcement Learning
- Utility-Optimized Local Differential Privacy Mechanisms for Distribution Estimation
- On the Gap Between Strict-Saddles and True Convexity: An Omega(log d) Lower Bound for Eigenvector Approximation
- A note on the approximate admissibility of regularized estimators in the Gaussian sequence model
- Optimal distributed composite testing in high-dimensional Gaussian models with 1-bit communication
- The Renyi Gaussian Process: Towards Improved Generalization