3 papers
cs.CL2026
PanCanBench: A Comprehensive Benchmark for Evaluating Large Language Models in Pancreatic Oncology
Yimin Zhao, Sheela R. Damle, Simone E. Dekker +13
Large language models (LLMs) have achieved expert-level performance on standardized examinations, yet multiple-choice accuracy poorly reflects real-world clinical utility and safet…
stat.ML2025
Do We Really Even Need Data? A Modern Look at Drawing Inference with Predicted Data
Stephen Salerno, Kentaro Hoffman, Awan Afiaz +3
As artificial intelligence and machine learning tools become more accessible, and scientists face new obstacles to data collection (e.g., rising costs, declining survey response ra…
stat.ME2024
ipd: An R Package for Conducting Inference on Predicted Data
Stephen Salerno, Jiacheng Miao, Awan Afiaz +5
Summary: ipd is an open-source R software package for the downstream modeling of an outcome and its associated features where a potentially sizable portion of the outcome data has…