Dangers of Bayesian Model Averaging under Covariate Shift
arXiv:2106.11905
Abstract
Approximate Bayesian inference for neural networks is considered a robust alternative to standard training, often providing good performance on out-of-distribution data. However, Bayesian neural networks (BNNs) with high-fidelity approximate inference via full-batch Hamiltonian Monte Carlo achieve poor generalization under covariate shift, even underperforming classical estimation. We explain this surprising result, showing how a Bayesian model average can in fact be problematic under covariate shift, particularly in cases where linear dependencies in the input features cause a lack of posterior contraction. We additionally show why the same issue does not affect many approximate inference procedures, or classical maximum a-posteriori (MAP) training. Finally, we propose novel priors that improve the robustness of BNNs to many sources of covariate shift.
NeurIPS 2021. Code is available at https://github.com/izmailovpavel/bnn_covariate_shift
References in corpus (10)
- ADADELTA: An Adaptive Learning Rate Method
- What Are Bayesian Neural Network Posteriors Really Like?
- Does Your Dermatology Classifier Know What It Doesn't Know? Detecting the Long-Tail of Unseen Conditions
- A Systematic Comparison of Bayesian Deep Learning Robustness in Diabetic Retinopathy Tasks
- Understanding the Failure Modes of Out-of-Distribution Generalization
- Out of Distribution Generalization in Machine Learning
- Bayesian Neural Network Priors Revisited
- Bayesian Inference with Certifiable Adversarial Robustness
- Ensemble Model Patching: A Parameter-Efficient Variational Bayesian Neural Network
- Loss Surface Simplexes for Mode Connecting Volumes and Fast Ensembling