3 papers
stat.ML2026
ERICA: Quantifying Replicability of Cluster Analysis
Siamak K. Sorooshyari, Manuel A. Rivas, Robert Tibshirani
Despite being ubiquitous in science, clustering lacks a unified framework for quantitatively evaluating the replicability of its results. We present evaluating replicability via it…
stat.ME2026
Univariate-Guided Sparse Regression for Biobank-Scale High-Dimensional Omics Data
Joshua Richland, Tuomo Kiiskinen, William Wang +5
We present a scalable framework for computing polygenic risk scores (PRS) in high-dimensional genomic settings using the recently introduced Univariate-Guided Sparse Regression (un…
stat.ME2024
Pretraining and the Lasso
Erin Craig, Mert Pilanci, Thomas Le Menestrel +7
Pretraining is a popular and powerful paradigm in machine learning to pass information from one model to another. As an example, suppose one has a modest-sized dataset of images of…