Covariate-Assisted Bayesian Graph Learning for Heterogeneous Data
arXiv:2308.07806 · doi:10.1080/01621459.2023.2233744
Abstract
In a traditional Gaussian graphical model, data homogeneity is routinely assumed with no extra variables affecting the conditional independence. In modern genomic datasets, there is an abundance of auxiliary information, which often gets under-utilized in determining the joint dependency structure. In this article, we consider a Bayesian approach to model undirected graphs underlying heterogeneous multivariate observations with additional assistance from covariates. Building on product partition models, we propose a novel covariate-dependent Gaussian graphical model that allows graphs to vary with covariates so that observations whose covariates are similar share a similar undirected graph. To efficiently embed Gaussian graphical models into our proposed framework, we explore both Gaussian likelihood and pseudo-likelihood functions. For Gaussian likelihood, a G-Wishart distribution is used as a natural conjugate prior, and for the pseudo-likelihood, a product of Gaussian-conditionals is used. Moreover, the proposed model has large prior support and is flexible to approximate any -Hölder conditional variance-covariance matrices with . We further show that based on the theory of fractional likelihood, the rate of posterior contraction is minimax optimal assuming the true density to be a Gaussian mixture with a known number of components. The efficacy of the approach is demonstrated via simulation studies and an analysis of a protein network for a breast cancer dataset assisted by mRNA gene expression as covariates.
58 pages, 12 figures, accepted by Journal of the American Statistical Association
References in corpus (6)
- Improved contact prediction in proteins: Using pseudolikelihoods to infer Potts models
- High-dimensional Ising model selection using -regularized logistic regression
- Bayesian variable selection with shrinking and diffusing priors
- A sparse conditional Gaussian graphical model for analysis of genetical genomics data
- Robust graphical modeling of gene networks using classical and alternative T-distributions
- Approximation of conditional densities by smooth mixtures of regressions