Conjugate-Computation Variational Inference : Converting Variational Inference in Non-Conjugate Models to Inferences in Conjugate Models
arXiv:1703.04265
Abstract
Variational inference is computationally challenging in models that contain both conjugate and non-conjugate terms. Methods specifically designed for conjugate models, even though computationally efficient, find it difficult to deal with non-conjugate terms. On the other hand, stochastic-gradient methods can handle the non-conjugate terms but they usually ignore the conjugate structure of the model which might result in slow convergence. In this paper, we propose a new algorithm called Conjugate-computation Variational Inference (CVI) which brings the best of the two worlds together -- it uses conjugate computations for the conjugate terms and employs stochastic gradients for the rest. We derive this algorithm by using a stochastic mirror-descent method in the mean-parameter space, and then expressing each gradient step as a variational inference in a conjugate model. We demonstrate our algorithm's applicability to a large class of models and establish its convergence. Our experimental results show that our method converges much faster than the methods that ignore the conjugate structure of the model.
Published in AI-Stats 2017. Fixed some typos. This version contains a short paragraph in the conclusions section which we could not add in the conference version due to space constraints
References in corpus (2)
Cited by in corpus (18)
- Practical Deep Learning with Bayesian Principles
- Fast and Scalable Bayesian Deep Learning by Weight-Perturbation in Adam
- Vprop: Variational Inference using RMSprop
- A Generalization Bound for Online Variational Inference
- Automatic structured variational inference
- Variational Adaptive-Newton Method for Explorative Learning
- The Bayesian Learning Rule
- Variational Inference In Pachinko Allocation Machines
- Provable Smoothness Guarantees for Black-Box Variational Inference
- Non-exponentially weighted aggregation: regret bounds for unbounded loss functions
- Stochastic Sequential Neural Networks with Structured Inference
- Flexible mean field variational inference using mixtures of non-overlapping exponential families
- BiSNN: Training Spiking Neural Networks with Binary Weights via Bayesian Learning
- Bayes-Newton Methods for Approximate Bayesian Inference with PSD Guarantees
- Practical Bayesian Learning of Neural Networks via Adaptive Optimisation Methods
- Black-box Optimizer with Implicit Natural Gradient
- Bayesian filtering unifies adaptive and non-adaptive neural network optimization methods
- Non-conjugate variational Bayes for pseudo-likelihood mixed effect models