3 papers
cs.CL2026
PanCanBench: A Comprehensive Benchmark for Evaluating Large Language Models in Pancreatic Oncology
Yimin Zhao, Sheela R. Damle, Simone E. Dekker +13
Large language models (LLMs) have achieved expert-level performance on standardized examinations, yet multiple-choice accuracy poorly reflects real-world clinical utility and safet…
stat.ML2025
Do We Really Even Need Data? A Modern Look at Drawing Inference with Predicted Data
Stephen Salerno, Kentaro Hoffman, Awan Afiaz +3
As artificial intelligence and machine learning tools become more accessible, and scientists face new obstacles to data collection (e.g., rising costs, declining survey response ra…
stat.ME2025
Some models are useful, but for how long?: A decision theoretic approach to choosing when to refit large-scale prediction models
Kentaro Hoffman, Stephen Salerno, Jeff Leek +1
Large-scale prediction models using tools from artificial intelligence (AI) or machine learning (ML) are increasingly common across a variety of industries and scientific domains.…