Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Ask-E: An Environment for Calibrated Question Generation
Sarah Pratt, Jae Sung Park, Scott Geng +1
Today, we improve models by training and evaluating them on problems at the frontier of their abilities. Creating such problems is itself a demanding task, requiring the ability to…
cs.CL2026
PanCanBench: A Comprehensive Benchmark for Evaluating Large Language Models in Pancreatic Oncology
Yimin Zhao, Sheela R. Damle, Simone E. Dekker +13
Large language models (LLMs) have achieved expert-level performance on standardized examinations, yet multiple-choice accuracy poorly reflects real-world clinical utility and safet…