activity
20242026
collaborators

5 papers

stat.ME2026

Univariate-Guided Sparse Regression for Biobank-Scale High-Dimensional Omics Data

Joshua Richland, Tuomo Kiiskinen, William Wang +5

We present a scalable framework for computing polygenic risk scores (PRS) in high-dimensional genomic settings using the recently introduced Univariate-Guided Sparse Regression (un…

stat.ME2025

Univariate-Guided Sparse Regression

Sourav Chatterjee, Trevor Hastie, Robert Tibshirani

In this paper, we introduce ``UniLasso'' -- a novel statistical method for sparse regression. This two-stage approach preserves the signs of the univariate coefficients and leverag…

stat.ME2025

Pre-validation Revisited

Jing Shang, Sourav Chatterjee, Trevor Hastie +1

Pre-validation is a way to build prediction model with two datasets of significantly different feature dimensions. Previous work showed that the asymptotic distribution of the resu…

q-bio.QM2024

MMIL: A novel algorithm for disease associated cell type discovery

Erin Craig, Timothy Keyes, Jolanda Sarno +5

Single-cell datasets often lack individual cell labels, making it challenging to identify cells associated with disease. To address this, we introduce Mixture Modeling for Multiple…

cs.LG2024

Using Pre-training and Interaction Modeling for ancestry-specific disease prediction in UK Biobank

Thomas Le Menestrel, Erin Craig, Robert Tibshirani +2

Recent genome-wide association studies (GWAS) have uncovered the genetic basis of complex traits, but show an under-representation of non-European descent individuals, underscoring…