Mitigating the Effects of Non-Identifiability on Inference for Bayesian Neural Networks with Latent Variables
arXiv:1911.00569
Abstract
Bayesian Neural Networks with Latent Variables (BNN+LVs) capture predictive uncertainty by explicitly modeling model uncertainty (via priors on network weights) and environmental stochasticity (via a latent input noise variable). In this work, we first show that BNN+LV suffers from a serious form of non-identifiability: explanatory power can be transferred between the model parameters and latent variables while fitting the data equally well. We demonstrate that as a result, in the limit of infinite data, the posterior mode over the network weights and latent variables is asymptotically biased away from the ground-truth. Due to this asymptotic bias, traditional inference methods may in practice yield parameters that generalize poorly and misestimate uncertainty. Next, we develop a novel inference procedure that explicitly mitigates the effects of likelihood non-identifiability during training and yields high-quality predictions as well as uncertainty estimates. We demonstrate that our inference method improves upon benchmark methods across a range of synthetic and real data-sets.
Accepted at JMLR 2022. Previously accepted at ICML's Uncertainty and Robustness in Deep Learning Workshop 2019
References in corpus (10)
- Weight Uncertainty in Neural Networks
- Probabilistic Backpropagation for Scalable Learning of Bayesian Neural Networks
- MINE: Mutual Information Neural Estimation
- Bayesian semi-supervised learning for uncertainty-calibrated prediction of molecular properties and active learning
- Lagging Inference Networks and Posterior Collapse in Variational Autoencoders
- Decomposition of Uncertainty in Bayesian Deep Learning for Efficient and Risk-sensitive Learning
- Model-Predictive Policy Learning with Uncertainty Regularization for Driving in Dense Traffic
- Gaussian Process Regression with Heteroscedastic or Non-Gaussian Residuals
- Avoiding Latent Variable Collapse With Generative Skip Models
- Variational Inference for Uncertainty on the Inputs of Gaussian Process Models