5 papers
Univariate-Guided Sparse Regression for Biobank-Scale High-Dimensional Omics Data
Joshua Richland, Tuomo Kiiskinen, William Wang +5
We present a scalable framework for computing polygenic risk scores (PRS) in high-dimensional genomic settings using the recently introduced Univariate-Guided Sparse Regression (un…
Univariate-Guided Sparse Regression
Sourav Chatterjee, Trevor Hastie, Robert Tibshirani
In this paper, we introduce ``UniLasso'' -- a novel statistical method for sparse regression. This two-stage approach preserves the signs of the univariate coefficients and leverag…
Pre-validation Revisited
Jing Shang, Sourav Chatterjee, Trevor Hastie +1
Pre-validation is a way to build prediction model with two datasets of significantly different feature dimensions. Previous work showed that the asymptotic distribution of the resu…
MMIL: A novel algorithm for disease associated cell type discovery
Erin Craig, Timothy Keyes, Jolanda Sarno +5
Single-cell datasets often lack individual cell labels, making it challenging to identify cells associated with disease. To address this, we introduce Mixture Modeling for Multiple…
Using Pre-training and Interaction Modeling for ancestry-specific disease prediction in UK Biobank
Thomas Le Menestrel, Erin Craig, Robert Tibshirani +2
Recent genome-wide association studies (GWAS) have uncovered the genetic basis of complex traits, but show an under-representation of non-European descent individuals, underscoring…