Noisy Natural Gradient as Variational Inference
arXiv:1712.02390
Abstract
Variational Bayesian neural nets combine the flexibility of deep learning with Bayesian uncertainty estimation. Unfortunately, there is a tradeoff between cheap but simple variational families (e.g.~fully factorized) or expensive and complicated inference procedures. We show that natural gradient ascent with adaptive weight noise implicitly fits a variational posterior to maximize the evidence lower bound (ELBO). This insight allows us to train full-covariance, fully factorized, or matrix-variate Gaussian variational posteriors using noisy versions of natural gradient, Adam, and K-FAC, respectively, making it possible to scale up to modern-size ConvNets. On standard regression benchmarks, our noisy K-FAC algorithm makes better predictions and matches Hamiltonian Monte Carlo's predictive variances better than existing methods. Its improved uncertainty estimates lead to more efficient exploration in active learning, and intrinsic motivation for reinforcement learning.
References in corpus (6)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- On Calibration of Modern Neural Networks
- Weight Uncertainty in Neural Networks
- Scalable trust-region method for deep reinforcement learning using Kronecker-factored approximation
- Vprop: Variational Inference using RMSprop
- Variational Adaptive-Newton Method for Explorative Learning
Cited by in corpus (20)
- A Simple Baseline for Bayesian Uncertainty in Deep Learning
- Task Agnostic Continual Learning Using Online Variational Bayes
- Fast and Scalable Bayesian Deep Learning by Weight-Perturbation in Adam
- Three Mechanisms of Weight Decay Regularization
- Understanding Short-Horizon Bias in Stochastic Meta-Optimization
- The Functional Neural Process
- The k-tied Normal Distribution: A Compact Parameterization of Gaussian Mean Field Posteriors in Bayesian Neural Networks
- Task Agnostic Continual Learning Using Online Variational Bayes with Fixed-Point Updates
- Global inducing point variational posteriors for Bayesian neural networks and deep Gaussian processes
- Delta-STN: Efficient Bilevel Optimization for Neural Networks using Structured Response Jacobians
- Eigenvalue Corrected Noisy Natural Gradient
- Prior choice affects ability of Bayesian neural networks to identify unknowns
- A statistical theory of cold posteriors in deep neural networks
- A Coordinate-Free Construction of Scalable Natural Gradient
- The Bayesian Learning Rule
- Hierarchical Gaussian Process Priors for Bayesian Neural Network Weights
- Network Automatic Pruning: Start NAP and Take a Nap
- Improving Classifier Confidence using Lossy Label-Invariant Transformations
- Variational Bayes Neural Network: Posterior Consistency, Classification Accuracy and Computational Challenges
- Stochastic Bayesian Neural Networks