Global inducing point variational posteriors for Bayesian neural networks and deep Gaussian processes
arXiv:2005.08140
Abstract
We consider the optimal approximate posterior over the top-layer weights in a Bayesian neural network for regression, and show that it exhibits strong dependencies on the lower-layer weights. We adapt this result to develop a correlated approximate posterior over the weights at all layers in a Bayesian neural network. We extend this approach to deep Gaussian processes, unifying inference in the two model classes. Our approximate posterior uses learned "global" inducing points, which are defined only at the input layer and propagated through the network to obtain inducing inputs at subsequent layers. By contrast, standard, "local", inducing point methods from the deep Gaussian process literature optimise a separate set of inducing inputs at every layer, and thus do not model correlations across layers. Our method gives state-of-the-art performance for a variational Bayesian method, without data augmentation or tempering, on CIFAR-10 of 86.7%, which is comparable to SGMCMC without tempering but with data augmentation (88% in Wenzel et al. 2020).
Accepted for publication at the 38th International Conference on Machine Learning (ICML 2021, PMLR 139), 33 pages
References in corpus (21)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- On Calibration of Modern Neural Networks
- Weight Uncertainty in Neural Networks
- Probabilistic Backpropagation for Scalable Learning of Bayesian Neural Networks
- A Comprehensive guide to Bayesian Convolutional Neural Network with Variational Inference
- Enhanced Convolutional Neural Tangent Kernels
- Quality of Uncertainty Quantification for Bayesian Neural Network Inference
- Convolutional Gaussian Processes
- Structured Stochastic Variational Inference
- Model Selection in Bayesian Neural Networks via Horseshoe Priors
- Efficient and Scalable Bayesian Neural Nets with Rank-1 Factors
- 'In-Between' Uncertainty in Bayesian Neural Networks
- Overpruning in Variational Bayesian Neural Networks
- Liberty or Depth: Deep Bayesian Neural Nets Do Not Need Complex Weight Posterior Approximations
- Bayesian Neural Network Priors Revisited
- A statistical theory of cold posteriors in deep neural networks
- Data augmentation in Bayesian neural networks and the cold posterior effect
- Hierarchical Gaussian Process Priors for Bayesian Neural Network Weights
- Importance Weighted Hierarchical Variational Inference
- Beyond the Mean-Field: Structured Deep Gaussian Processes Improve the Predictive Uncertainties
- Deep kernel processes
Cited by in corpus (8)
- A Review of Uncertainty Quantification in Deep Learning: Techniques, Applications and Challenges
- Finite Versus Infinite Neural Networks: an Empirical Study
- Bayesian Neural Network Priors Revisited
- A statistical theory of cold posteriors in deep neural networks
- Deep kernel processes
- Variational Laplace for Bayesian neural networks
- A variational approximate posterior for the deep Wishart process
- Conditional Deep Gaussian Processes: empirical Bayes hyperdata learning