Bayes-Newton Methods for Approximate Bayesian Inference with PSD Guarantees
arXiv:2111.01721
Abstract
We formulate natural gradient variational inference (VI), expectation propagation (EP), and posterior linearisation (PL) as extensions of Newton's method for optimising the parameters of a Bayesian posterior distribution. This viewpoint explicitly casts inference algorithms under the framework of numerical optimisation. We show that common approximations to Newton's method from the optimisation literature, namely Gauss-Newton and quasi-Newton methods (e.g., the BFGS algorithm), are still valid under this 'Bayes-Newton' framework. This leads to a suite of novel algorithms which are guaranteed to result in positive semi-definite (PSD) covariance matrices, unlike standard VI and EP. Our unifying viewpoint provides new insights into the connections between various inference schemes. All the presented methods apply to any model with a Gaussian prior and non-conjugate likelihood, which we demonstrate with (sparse) Gaussian processes and state space models.
Code for methods and experiments: https://github.com/AaltoML/BayesNewton
References in corpus (8)
- Scalable Variational Gaussian Process Classification
- Conjugate-Computation Variational Inference : Converting Variational Inference in Non-Conjugate Models to Inferences in Conjugate Models
- Fast Convergent Algorithms for Expectation Propagation Approximate Bayesian Inference
- Partitioned Variational Inference: A unified framework encompassing federated and continual learning
- State Space Gaussian Processes with Non-Gaussian Likelihood
- Approximate Inference Turns Deep Networks into Gaussian Processes
- The Bayesian Learning Rule
- State Space Expectation Propagation: Efficient Inference Schemes for Temporal Gaussian Processes