2 papers
cs.CL2026
ER-Reason: A Benchmark Dataset for LLM Clinical Reasoning in the Emergency Room
Nikita Mehandru, Niloufar Golchini, Namrata Garg +9
Existing benchmarks for evaluating the clinical reasoning capabilities of large language models (LLMs) often lack a clear definition of "clinical reasoning" as a construct, fail to…
cs.CL2025
Medical Large Language Model Benchmarks Should Prioritize Construct Validity
Ahmed Alaa, Thomas Hartvigsen, Niloufar Golchini +4
Medical large language models (LLMs) research often makes bold claims, from encoding clinical knowledge to reasoning like a physician. These claims are usually backed by evaluation…