Yimin Zhao, Sheela R. Damle, Simone E. Dekker +13
Large language models (LLMs) have achieved expert-level performance on standardized examinations, yet multiple-choice accuracy poorly reflects real-world clinical utility and safet…