Data augmentation in Bayesian neural networks and the cold posterior effect
arXiv:2106.05586
Abstract
Bayesian neural networks that incorporate data augmentation implicitly use a ``randomly perturbed log-likelihood [which] does not have a clean interpretation as a valid likelihood function'' (Izmailov et al. 2021). Here, we provide several approaches to developing principled Bayesian neural networks incorporating data augmentation. We introduce a ``finite orbit'' setting which allows likelihoods to be computed exactly, and give tight multi-sample bounds in the more usual ``full orbit'' setting. These models cast light on the origin of the cold posterior effect. In particular, we find that the cold posterior effect persists even in these principled models incorporating data augmentation. This suggests that the cold posterior effect cannot be dismissed as an artifact of data augmentation using incorrect likelihoods.
References in corpus (12)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Fashion-MNIST: a Novel Image Dataset for Benchmarking Machine Learning Algorithms
- Weight Uncertainty in Neural Networks
- What Are Bayesian Neural Network Posteriors Really Like?
- Augment your batch: better training with larger batches
- MultiGrain: a unified image embedding for classes and instances
- Learning Invariances in Neural Networks
- Bayesian Neural Network Priors Revisited
- BNNpriors: A library for Bayesian neural network inference with different prior distributions
- Improving Transformation Invariance in Contrastive Representation Learning
- A statistical theory of cold posteriors in deep neural networks
- Drawing Multiple Augmentation Samples Per Image During Training Efficiently Decreases Test Error