2 papers
cs.CL2025
MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Suhana Bedi, Hejie Cui, Miguel Fuentes +78
While large language models (LLMs) achieve near-perfect scores on medical licensing exams, these evaluations inadequately reflect the complexity and diversity of real-world clinica…
cs.AI2025
Superhuman performance of a large language model on the reasoning tasks of a physician
Peter G. Brodeur, Thomas A. Buckley, Zahir Kanjee +22
A seminal paper published by Ledley and Lusted in 1959 introduced complex clinical diagnostic reasoning cases as the gold standard for the evaluation of expert medical computing sy…