Bayes meets Bernstein at the Meta Level: an Analysis of Fast Rates in Meta-Learning with PAC-Bayes
arXiv:2302.11709
Abstract
Bernstein's condition is a key assumption that guarantees fast rates in machine learning. For example, the Gibbs algorithm with prior has an excess risk in , as opposed to the standard , where denotes the number of observations and is a complexity parameter which depends on the prior . In this paper, we examine the Gibbs algorithm in the context of meta-learning, i.e., when learning the prior from tasks (with observations each) generated by a meta distribution. Our main result is that Bernstein's condition always holds at the meta level, regardless of its validity at the observation level. This implies that the additional cost to learn the Gibbs prior , which will reduce the term across tasks, is in , instead of the expected . We further illustrate how this result improves on standard rates in three different settings: discrete priors, Gaussian priors and mixture of Gaussians priors.
References in corpus (9)
- On the properties of variational approximations of Gibbs posteriors
- Provable Guarantees for Gradient-Based Meta-Learning
- Generalization Bounds For Meta-Learning: An Information-Theoretic Analysis
- Bridging the Gap Between Practice and PAC-Bayes Theory in Few-Shot Meta-Learning
- Towards a Unified Information-Theoretic Framework for Generalization
- Evaluated CMI Bounds for Meta Learning: Tightness and Expressiveness
- PAC-Bayes Bounds for Meta-learning with Data-Dependent Prior
- Scalable PAC-Bayesian Meta-Learning via the PAC-Optimal Hyper-Posterior: From Theory to Practice
- Is Bayesian Model-Agnostic Meta Learning Better than Model-Agnostic Meta Learning, Provably?