On the properties of variational approximations of Gibbs posteriors
arXiv:1506.04091
Abstract
The PAC-Bayesian approach is a powerful set of techniques to derive non- asymptotic risk bounds for random estimators. The corresponding optimal distribution of estimators, usually called the Gibbs posterior, is unfortunately intractable. One may sample from it using Markov chain Monte Carlo, but this is often too slow for big datasets. We consider instead variational approximations of the Gibbs posterior, which are fast to compute. We undertake a general study of the properties of such approximations. Our main finding is that such a variational approximation has often the same rate of convergence as the original PAC-Bayesian procedure it approximates. We specialise our results to several learning tasks (classification, ranking, matrix completion),discuss how to implement a variational approximation in each case, and illustrate the good properties of said approximation on real datasets.
References in corpus (4)
Cited by in corpus (21)
- User-friendly introduction to PAC-Bayes bounds
- Tighter risk certificates for neural networks
- Gibbs posterior concentration rates under sub-exponential type losses
- Normalized Flat Minima: Exploring Scale Invariant Definition of Flat Minima for Neural Networks using PAC-Bayesian Analysis
- Empirical Risk Minimization with Relative Entropy Regularization
- Bayesian Model Selection via Mean-Field Variational Approximation
- From bilinear regression to inductive matrix completion: a quasi-Bayesian analysis
- Meta-Learning PAC-Bayes Priors in Model Averaging
- High-dimensional sparse classification using exponential weighting with empirical hinge loss
- Adaptive variational Bayes: Optimality, computation and applications
- PAC-Bayes Generalisation Bounds for Dynamical Systems Including Stable RNNs
- An efficient adaptive MCMC algorithm for Pseudo-Bayesian quantum tomography
- PAC-Bayes unleashed: generalisation bounds with unbounded losses
- Bayes meets Bernstein at the Meta Level: an Analysis of Fast Rates in Meta-Learning with PAC-Bayes
- Wasserstein PAC-Bayes Learning: Exploiting Optimisation Guarantees to Explain Generalisation
- Semiparametric inference using fractional posteriors
- Learning via Wasserstein-Based High Probability Generalisation Bounds
- Explicit construction of the minimum error variance estimator for stochastic LTI state-space systems
- PAC-Bayesian bounds for learning LTI-ss systems with input from empirical loss
- PAC-Bayes-Chernoff bounds for unbounded losses
- A note on regularised NTK dynamics with an application to PAC-Bayesian training