A PAC-Bayesian Tutorial with A Dropout Bound
arXiv:1307.2118
Abstract
This tutorial gives a concise overview of existing PAC-Bayesian theory focusing on three generalization bounds. The first is an Occam bound which handles rules with finite precision parameters and which states that generalization loss is near training loss when the number of bits needed to write the rule is small compared to the sample size. The second is a PAC-Bayesian bound providing a generalization guarantee for posterior distributions rather than for individual rules. The PAC-Bayesian bound naturally handles infinite precision rule parameters, regularization, {\em provides a bound for dropout training}, and defines a natural notion of a single distinguished PAC-Bayesian posterior distribution. The third bound is a training-variance bound --- a kind of bias-variance analysis but with bias replaced by expected training loss. The training-variance bound dominates the other bounds but is more difficult to interpret. It seems to suggest variance reduction methods such as bagging and may ultimately provide a more meaningful analysis of dropouts.
Cited by in corpus (33)
- Non-Vacuous Generalization Bounds at the ImageNet Scale: A PAC-Bayesian Compression Approach
- Risk Bounds for the Majority Vote: From a PAC-Bayesian Analysis to a Learning Algorithm
- User-friendly introduction to PAC-Bayes bounds
- Where is the Information in a Deep Neural Network?
- Tighter risk certificates for neural networks
- A New PAC-Bayesian Perspective on Domain Adaptation
- How much does your data exploration overfit? Controlling bias via information usage
- To Drop or Not to Drop: Robustness, Consistency and Differential Privacy Properties of Dropout
- Altitude Training: Strong Bounds for Single-Layer Dropout
- PAC-Bayes under potentially heavy tails
- PAC-Bayes and Domain Adaptation
- PAC-Bayes Mini-tutorial: A Continuous Union Bound
- Information Complexity and Generalization Bounds
- On Convergence and Generalization of Dropout Training
- Dropout: Explicit Forms and Capacity Control
- Generalization bounds for deep learning
- Adversarial Robustness Guarantees for Classification with Gaussian Processes
- Information-Theoretic Generalization Bounds for Stochastic Gradient Descent
- Dynamics and Reachability of Learning Tasks
- PAC-Bayes Analysis Beyond the Usual Bounds
- Calibrating Noise to Variance in Adaptive Data Analysis
- Agnostic insurability of model classes
- Marginalizing Corrupted Features
- Dropout Rademacher Complexity of Deep Neural Networks
- Bregman Divergence Bounds and Universality Properties of the Logarithmic Loss
- How Tight Can PAC-Bayes be in the Small Data Regime?
- What training reveals about neural network complexity
- A Free-Energy Principle for Representation Learning
- Structured Dropout Variational Inference for Bayesian Neural Networks
- Optimal Posteriors for Chi-squared Divergence based PAC-Bayesian Bounds and Comparison with KL-divergence based Optimal Posteriors and Cross-Validation Procedure
- Entropy-SGD optimizes the prior of a PAC-Bayes bound: Generalization properties of Entropy-SGD and data-dependent priors
- On the generalization of bayesian deep nets for multi-class classification
- Generalization Bounds for Meta-Learning via PAC-Bayes and Uniform Stability