Bayesian Hypernetworks
arXiv:1710.04759
Abstract
We study Bayesian hypernetworks: a framework for approximate Bayesian inference in neural networks. A Bayesian hypernetwork $\h$ is a neural network which learns to transform a simple noise distribution, $p(\vecε) = \N(\vec 0,\mat I)$, to a distribution $q(\pp) := q(h(\vecε))$ over the parameters $\pp$ of another neural network (the "primary network")\@. We train with variational inference, using an invertible $\h$ to enable efficient estimation of the variational lower bound on the posterior $p(\pp | \D)$ via sampling. In contrast to most methods for Bayesian deep learning, Bayesian hypernets can represent a complex multimodal approximate posterior with correlations between parameters, while enabling cheap iid sampling of~$q(\pp)$. In practice, Bayesian hypernets can provide a better defense against adversarial examples than dropout, and also exhibit competitive performance on a suite of tasks which evaluate model uncertainty, including regularization, active learning, and anomaly detection.
David Krueger and Chin-Wei Huang contributed equally
References in corpus (6)
- Variational Inference with Normalizing Flows
- Probabilistic Backpropagation for Scalable Learning of Bayesian Neural Networks
- Markov Chain Monte Carlo and Variational Inference: Bridging the Gap
- MADE: Masked Autoencoder for Distribution Estimation
- Arbitrary Style Transfer in Real-time with Adaptive Instance Normalization
- Learning Visual Reasoning Without Strong Priors
Cited by in corpus (11)
- A Review of Uncertainty Quantification in Deep Learning: Techniques, Applications and Challenges
- Subspace Inference for Bayesian Deep Learning
- Infinitely Deep Bayesian Neural Networks with Stochastic Differential Equations
- Predictive Uncertainty Quantification with Compound Density Networks
- Regularization-Agnostic Compressed Sensing MRI Reconstruction with Hypernetworks
- VFunc: a Deep Generative Model for Functions
- Stochastic Neural Network with Kronecker Flow
- Posterior Meta-Replay for Continual Learning
- Variational Hyper RNN for Sequence Modeling
- Are Bayesian neural networks intrinsically good at out-of-distribution detection?
- Alpha-Divergences in Variational Dropout