Prior-preconditioned conjugate gradient method for accelerated Gibbs sampling in "large & large " Bayesian sparse regression
arXiv:1810.12437 · doi:10.1080/01621459.2022.2057859
Abstract
In a modern observational study based on healthcare databases, the number of observations and of predictors typically range in the order of ~ and of ~ . Despite the large sample size, data rarely provide sufficient information to reliably estimate such a large number of parameters. Sparse regression techniques provide potential solutions, one notable approach being the Bayesian methods based on shrinkage priors. In the "large n & large p" setting, however, posterior computation encounters a major bottleneck at repeated sampling from a high-dimensional Gaussian distribution, whose precision matrix is expensive to compute and factorize. In this article, we present a novel algorithm to speed up this bottleneck based on the following observation: we can cheaply generate a random vector such that the solution to the linear system has the desired Gaussian distribution. We can then solve the linear system by the conjugate gradient (CG) algorithm through matrix-vector multiplications by ; this involves no explicit factorization or calculation of itself. Rapid convergence of CG in this context is guaranteed by the theory of prior-preconditioning we develop. We apply our algorithm to a clinically relevant large-scale observational study with n = 72,489 patients and p = 22,175 clinical covariates, designed to assess the relative risk of adverse events from two alternative blood anti-coagulants. Our algorithm demonstrates an order of magnitude speed-up in posterior inference, in our case cutting the computation time from two weeks to less than a day.
36 pages, 7 figures + Supplement; Software package available --- see documentation at https://bayes-bridge.readthedocs.io and source code at https://github.com/aki-nishimura/bayes-bridge
References in corpus (6)
- Matching Methods for Causal Inference: A Review and a Look Forward
- Sparsity information and regularization in the horseshoe and other shrinkage priors
- Practical Bayesian Modeling and Inference for Massive Spatial Datasets On Modest Computing Environments
- Bayes Shrinkage at GWAS scale: Convergence and Approximation Theory of a Scalable MCMC Algorithm for the Horseshoe Prior
- Fast model-fitting of Bayesian variable selection regression using the iterative complex factorization algorithm
- A systematic approach to improving the reliability and scale of evidence from health care data