Fast Rates for General Unbounded Loss Functions: from ERM to Generalized Bayes
arXiv:1605.00252
Abstract
We present new excess risk bounds for general unbounded loss functions including log loss and squared loss, where the distribution of the losses may be heavy-tailed. The bounds hold for general estimators, but they are optimized when applied to -generalized Bayesian, MDL, and empirical risk minimization estimators. In the case of log loss, the bounds imply convergence rates for generalized Bayesian inference under misspecification in terms of a generalization of the Hellinger metric as long as the learning rate is set correctly. For general loss functions, our bounds rely on two separate conditions: the -GRIP (generalized reversed information projection) conditions, which control the lower tail of the excess loss; and the newly introduced witness condition, which controls the upper tail. The parameter in the -GRIP conditions determines the achievable rate and is akin to the exponent in the Tsybakov margin condition and the Bernstein condition for bounded losses, which the -GRIP conditions generalize; favorable in combination with small model complexity leads to rates. The witness condition allows us to connect the excess risk to an "annealed" version thereof, by which we generalize several previous results connecting Hellinger and Rényi divergence to KL divergence.
accepted to JMLR pending minor final modifications
Cited by in corpus (21)
- The no-free-lunch theorems of supervised learning
- User-friendly introduction to PAC-Bayes bounds
- A Primer on PAC-Bayesian Learning
- A comparison of learning rate selection methods in generalized Bayesian inference
- Safe-Bayesian Generalized Linear Regression
- Learning with Non-Convex Truncated Losses by SGD
- Information Complexity and Generalization Bounds
- Data-dependent PAC-Bayes priors via differential privacy
- PAC-Bayes, MAC-Bayes and Conditional Mutual Information: Fast rate bounds that handle general VC classes
- Minimax Rates for Conditional Density Estimation via Empirical Entropy
- On Empirical Risk Minimization with Dependent and Heavy-Tailed Data
- Bayes meets Bernstein at the Meta Level: an Analysis of Fast Rates in Meta-Learning with PAC-Bayes
- Calibrating generalized predictive distributions
- Semiparametric inference using fractional posteriors
- Computationally efficient variational-like approximations of possibilistic inferential models
- Sequential prediction under log-loss and misspecification
- Sharper convergence bounds of Monte Carlo Rademacher Averages through Self-Bounding functions
- A Note on High-Probability versus In-Expectation Guarantees of Generalization Bounds in Machine Learning
- Generalized Bayes Approach to Inverse Problems with Model Misspecification
- Unifying Variational Inference and PAC-Bayes for Supervised Learning that Scales
- Minimax Optimal Quantile and Semi-Adversarial Regret via Root-Logarithmic Regularizers