User-friendly introduction to PAC-Bayes bounds
arXiv:2110.11216 · doi:10.1561/2200000100
Abstract
Aggregated predictors are obtained by making a set of basic predictors vote according to some weights, that is, to some probability distribution. Randomized predictors are obtained by sampling in a set of basic predictors, according to some prescribed probability distribution. Thus, aggregated and randomized predictors have in common that they are not defined by a minimization problem, but by a probability distribution on the set of predictors. In statistical learning theory, there is a set of tools designed to understand the generalization ability of such procedures: PAC-Bayesian or PAC-Bayes bounds. Since the original PAC-Bayes bounds of D. McAllester, these tools have been considerably improved in many directions (we will for example describe a simplified version of the localization technique of O. Catoni that was missed by the community, and later rediscovered as "mutual information bounds"). Very recently, PAC-Bayes bounds received a considerable attention: for example there was workshop on PAC-Bayes at NIPS 2017, "(Almost) 50 Shades of Bayesian Learning: PAC-Bayesian trends and insights", organized by B. Guedj, F. Bach and P. Germain. One of the reason of this recent success is the successful application of these bounds to neural networks by G. Dziugaite and D. Roy. An elementary introduction to PAC-Bayes theory is still missing. This is an attempt to provide such an introduction.
References in corpus (58)
- A PAC-Bayesian Approach to Spectrally-Normalized Margin Bounds for Neural Networks
- Information-theoretic analysis of generalization capability of learning algorithms
- A PAC-Bayesian bound for Lifelong Learning
- Sparse Regression Learning by Aggregation and Langevin Monte-Carlo
- Online Learning: A Modern Introduction Using Convex Optimization
- Gibbs posterior for variable selection in high-dimensional classification and data mining
- Sharpness-Aware Minimization for Efficiently Improving Generalization
- A Note on the PAC Bayesian Theorem
- Meta-Learning by Adjusting Priors Based on Extended PAC-Bayes Theory
- Calibrating general posterior credible regions
- Non-Vacuous Generalization Bounds at the ImageNet Scale: A PAC-Bayesian Compression Approach
- Fast learning rates in statistical inference through aggregation
- On the properties of variational approximations of Gibbs posteriors
- A PAC-Bayesian Tutorial with A Dropout Bound
- The Many Faces of Exponential Weights in Online Learning
- Generalized Variational Inference: Three arguments for deriving new Posteriors
- Fast Rates for General Unbounded Loss Functions: from ERM to Generalized Bayes
- Sparse single-index model
- PAC-Bayesian Bounds for Randomized Empirical Risk Minimizers
- Optimal rates and adaptation in the single-index model using aggregation
- Consistency of Variational Bayes Inference for Estimation and Model Selection in Mixtures
- A PAC-Bayesian Analysis of Randomized Learning with Application to Stochastic Gradient Descent
- Generalization Bounds via Information Density and Conditional Information Density
- Learners that Use Little Information
- Simpler PAC-Bayesian Bounds for Hostile Data
- 1-bit Matrix Completion: PAC-Bayesian Analysis of a Variational Approximation
- Tighter risk certificates for neural networks
- Chromatic PAC-Bayes Bounds for Non-IID Data: Applications to Ranking and Stationary -Mixing Processes
- Bayesian methods for low-rank matrix estimation: short survey and theoretical study
- A New PAC-Bayesian Perspective on Domain Adaptation
- Gibbs posterior concentration rates under sub-exponential type losses
- Tighter expected generalization error bounds via Wasserstein distance
- PAC-Bayes under potentially heavy tails
- Dimension-free PAC-Bayesian bounds for the estimation of the mean of a random vector
- Practical bounds on the error of Bayesian posterior approximations: A nonasymptotic approach
- PAC-Bayes with Backprop
- Information Complexity and Generalization Bounds
- PAC-Bayes Mini-tutorial: A Continuous Union Bound
- On the Generalization Gap in Reparameterizable Reinforcement Learning
- PAC-Bayes-Bernstein Inequality for Martingales and its Application to Multiarmed Bandits
- PAC-Bayesian AUC classification and scoring
- Exponential weights in multivariate regression and a low-rankness favoring prior
- PAC-Bayes Information Bottleneck
- Fast-rate PAC-Bayes Generalization Bounds via Shifted Rademacher Processes
- Information-Theoretic Generalization Bounds for Stochastic Gradient Descent
- The Bayesian Learning Rule
- PAC-Bayes, MAC-Bayes and Conditional Mutual Information: Fast rate bounds that handle general VC classes
- PAC-Bayes Bounds for Meta-learning with Data-Dependent Prior
- PAC-Bayesian Contrastive Unsupervised Representation Learning
- Asymptotic Consistency of Rényi-Approximate Posteriors
- Upper and Lower Bounds on the Performance of Kernel PCA
- On the Robustness to Misspecification of -Posteriors and Their Variational Approximations
- On the Difficulty of Unbiased Alpha Divergence Minimization
- Characterizing the Generalization Error of Gibbs Algorithm with Symmetrized KL information
- PAC-Bayes Analysis of Sentence Representation
- PAC-Bayesian Generalization Bounds for MultiLayer Perceptrons
- New bounds for -means and information -means
- How Tight Can PAC-Bayes be in the Small Data Regime?
Cited by in corpus (8)
- Generalization Bounds for Neural Belief Propagation Decoders
- High-dimensional sparse classification using exponential weighting with empirical hinge loss
- Misclassification bounds for PAC-Bayesian sparse deep learning
- Progress in Self-Certified Neural Networks
- A PAC-Bayesian Framework for Optimal Control with Stability Guarantees
- Comparing Comparators in Generalization Bounds
- PAC-Bayes Bounds for High-Dimensional Multi-Index Models with Unknown Active Dimension
- A PAC-Bayes oracle inequality for sparse neural networks