Bayesian variable selection regression for genome-wide association studies and other large-scale problems
arXiv:1110.6019 · doi:10.1214/11-AOAS455
Abstract
We consider applying Bayesian Variable Selection Regression, or BVSR, to genome-wide association studies and similar large-scale regression problems. Currently, typical genome-wide association studies measure hundreds of thousands, or millions, of genetic variants (SNPs), in thousands or tens of thousands of individuals, and attempt to identify regions harboring SNPs that affect some phenotype or outcome of interest. This goal can naturally be cast as a variable selection regression problem, with the SNPs as the covariates in the regression. Characteristic features of genome-wide association studies include the following: (i) a focus primarily on identifying relevant variables, rather than on prediction; and (ii) many relevant covariates may have tiny effects, making it effectively impossible to confidently identify the complete "correct" subset of variables. Taken together, these factors put a premium on having interpretable measures of confidence for individual covariates being included in the model, which we argue is a strength of BVSR compared with alternatives such as penalized regression methods. Here we focus primarily on analysis of quantitative phenotypes, and on appropriate prior specification for BVSR in this setting, emphasizing the idea of considering what the priors imply about the total proportion of variance in outcome explained by relevant covariates. We also emphasize the potential for BVSR to estimate this proportion of variance explained, and hence shed light on the issue of "missing heritability" in genome-wide association studies.
Published in at http://dx.doi.org/10.1214/11-AOAS455 the Annals of Applied Statistics (http://www.imstat.org/aoas/) by the Institute of Mathematical Statistics (http://www.imstat.org)
References in corpus (1)
Cited by in corpus (15)
- Bayesian model averaging: A systematic review and conceptual classification
- Gene Hunting with Knockoffs for Hidden Markov Models
- Incorporating biological information into linear models: A Bayesian approach to the selection of pathways and genes
- Genome scans for detecting footprints of local adaptation using a Bayesian factor model
- Variable Selection for Nonparametric Gaussian Process Priors: Models and Computational Strategies
- Bayesian Model Selection in Complex Linear Systems, as Illustrated in Genetic Association Studies
- Detection boundary and Higher Criticism approach for rare and weak genetic effects
- Efficient inference for genetic association studies with multiple outcomes
- A novel algorithm for simultaneous SNP selection in high-dimensional genome-wide association studies
- Sticky PDMP samplers for sparse and local inference problems
- A Scalable Empirical Bayes Approach to Variable Selection in Generalized Linear Models
- Variable Selection with ABC Bayesian Forests
- AcSel: selecting variables with accuracy in correlated datasets
- selectBoost: a general algorithm to enhance the performance of variable selection methods in correlated datasets
- Genetic variant selection: learning across traits and sites