3 papers
cs.CL2026
CLExEval: A Human-in-the-Loop Framework for Qualitative Evaluation of LLM Clinical Reasoning
Ajmal M., Abin Roy, Afthab Salam Kanniyan +4
Large Language Models (LLMs) achieve strong results on many medical benchmarks, but their clinical reasoning remains difficult to evaluate reliably. A central risk is an evaluation…
cs.CL2026
ER-Reason: A Benchmark Dataset for LLM Clinical Reasoning in the Emergency Room
Nikita Mehandru, Niloufar Golchini, Namrata Garg +9
Existing benchmarks for evaluating the clinical reasoning capabilities of large language models (LLMs) often lack a clear definition of "clinical reasoning" as a construct, fail to…
cs.CL2025
Asking the Right Questions: Benchmarking Large Language Models in the Development of Clinical Consultation Templates
Liam G. McCoy, Fateme Nateghi Haredasht, Kanav Chopra +15
This study evaluates the capacity of large language models (LLMs) to generate structured clinical consultation templates for electronic consultation. Using 145 expert-crafted templ…