Stochastic Quasi-Newton Langevin Monte Carlo
arXiv:1602.03442
Abstract
Recently, Stochastic Gradient Markov Chain Monte Carlo (SG-MCMC) methods have been proposed for scaling up Monte Carlo computations to large data problems. Whilst these approaches have proven useful in many applications, vanilla SG-MCMC might suffer from poor mixing rates when random variables exhibit strong couplings under the target densities or big scale differences. In this study, we propose a novel SG-MCMC method that takes the local geometry into account by using ideas from Quasi-Newton optimization methods. These second order methods directly approximate the inverse Hessian by using a limited history of samples and their gradients. Our method uses dense approximations of the inverse Hessian while keeping the time and memory complexities linear with the dimension of the problem. We provide a formal theoretical analysis where we show that the proposed method is asymptotically unbiased and consistent with the posterior expectations. We illustrate the effectiveness of the approach on both synthetic and real datasets. Our experiments on two challenging applications show that our method achieves fast convergence rates similar to Riemannian approaches while at the same time having low computational requirements similar to diagonal preconditioning approaches.
Published in ICML 2016, International Conference on Machine Learning 2016, New York, NY, USA
References in corpus (8)
- A Complete Recipe for Stochastic Gradient MCMC
- Preconditioned Stochastic Gradient Langevin Dynamics for Deep Neural Networks
- On the Convergence of Stochastic Gradient MCMC Algorithms with High-Order Integrators
- Privacy for Free: Posterior Sampling and Stochastic Gradient Monte Carlo
- Consistency and fluctuations for stochastic gradient Langevin dynamics
- Parallel Stochastic Gradient Markov Chain Monte Carlo for Matrix Factorisation Models
- Large-Scale Distributed Bayesian Matrix Factorization using Stochastic Gradient MCMC
- HAMSI: A Parallel Incremental Optimization Algorithm Using Quadratic Approximations for Solving Partially Separable Problems
Cited by in corpus (9)
- Deep Latent Dirichlet Allocation with Topic-Layer-Adaptive Stochastic Gradient Riemannian MCMC
- Global Convergence of Stochastic Gradient Hamiltonian Monte Carlo for Non-Convex Stochastic Optimization: Non-Asymptotic Performance Bounds and Momentum-Based Acceleration
- Bayesian Pose Graph Optimization via Bingham Distributions and Tempered Geodesic MCMC
- Bayesian Sparse learning with preconditioned stochastic gradient MCMC and its applications
- A Convergence Analysis for A Class of Practical Variance-Reduction Stochastic Gradient MCMC
- Efficient constrained sampling via the mirror-Langevin algorithm
- Guaranteed inference in topic models
- Stochastic quasi-Newton with adaptive step lengths for large-scale problems
- Mirrored Langevin Dynamics