activity
20242026
collaborators

5 papers

stat.ML2026

Spatially Robust Inference with Predicted and Missing at Random Labels

Stephen Salerno, Zhenke Wu, Tyler McCormick

When outcome data are expensive or onerous to collect, scientists increasingly substitute predictions from machine learning and AI models for unlabeled cases, a process which has c…

stat.ML2025

Do We Really Even Need Data? A Modern Look at Drawing Inference with Predicted Data

Stephen Salerno, Kentaro Hoffman, Awan Afiaz +3

As artificial intelligence and machine learning tools become more accessible, and scientists face new obstacles to data collection (e.g., rising costs, declining survey response ra…

stat.ME2025

What's the Weight? Estimating Controlled Outcome Differences in Complex Surveys for Health Disparities Research

Stephen Salerno, Emily K. Roberts, Belinda L. Needham +4

In this work, we are motivated by the problem of estimating racial disparities in health outcomes, specifically the average controlled difference (ACD) in telomere length between B…

stat.ME2025

Some models are useful, but for how long?: A decision theoretic approach to choosing when to refit large-scale prediction models

Kentaro Hoffman, Stephen Salerno, Jeff Leek +1

Large-scale prediction models using tools from artificial intelligence (AI) or machine learning (ML) are increasingly common across a variety of industries and scientific domains.…

stat.ME2024

ipd: An R Package for Conducting Inference on Predicted Data

Stephen Salerno, Jiacheng Miao, Awan Afiaz +5

Summary: ipd is an open-source R software package for the downstream modeling of an outcome and its associated features where a potentially sizable portion of the outcome data has…