paper

Cross-trait prediction accuracy of high-dimensional ridge-type estimators in genome-wide association studies

arXiv:1911.10142

Abstract

Marginal association summary statistics have attracted great attention in statistical genetics, mainly because the primary results of most genome-wide association studies (GWAS) are produced by marginal screening. In this paper, we study the prediction accuracy of marginal estimator in dense (or sparsity free) high-dimensional settings with , , and . We consider a general correlation structure among the features and allow an unknown subset of them to be signals. As the marginal estimator can be viewed as a ridge estimator with regularization parameter , we further investigate a class of ridge-type estimators in a unifying framework, including the popular best linear unbiased prediction (BLUP) in genetics. We find that the influence of on out-of-sample prediction accuracy heavily depends on . Though selecting an optimal can be important when and are comparable, it turns out that the out-of-sample of ridge-type estimators becomes near-optimal for any as increases. For example, when features are independent, the out-of-sample is always bounded by from above and is largely invariant to given large (say, ). We also find that in-sample has completely different patterns and depends much more on than out-of-sample . In practice, our analysis delivers useful messages for genome-wide polygenic risk prediction and computation-accuracy trade-off in dense high-dimensions. We numerically illustrate our results in simulation studies and a real data example.

References in corpus (3)

Cited by in corpus (1)