activity
20242026
most citedBRIDGE: Benchmarking Large Language Models for Understanding Real-world Clinical Practice Text

3 citations · 6 across the 8 of their papers we have counts for

collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL2026

CLExEval: A Human-in-the-Loop Framework for Qualitative Evaluation of LLM Clinical Reasoning

Ajmal M., Abin Roy, Afthab Salam Kanniyan +4

Large Language Models (LLMs) achieve strong results on many medical benchmarks, but their clinical reasoning remains difficult to evaluate reliably. A central risk is an evaluation…

cs.CL2026

ER-Reason: A Benchmark Dataset for LLM Clinical Reasoning in the Emergency Room

Nikita Mehandru, Niloufar Golchini, Namrata Garg +9

Existing benchmarks for evaluating the clinical reasoning capabilities of large language models (LLMs) often lack a clear definition of "clinical reasoning" as a construct, fail to…

cs.CL20263 cited

BRIDGE: Benchmarking Large Language Models for Understanding Real-world Clinical Practice Text

Jiageng Wu, Bowen Gu, Ren Zhou +14

Large language models (LLMs) hold great promise for medical applications and are evolving rapidly, with new models being released at an accelerated pace. However, benchmarking on l…

cs.CL2025

Advancing Conversational Diagnostic AI with Multimodal Reasoning

Khaled Saab, Jan Freyberg, Chunjong Park +33

Large Language Models (LLMs) have demonstrated great potential for conducting diagnostic conversations but evaluation has been largely limited to language-only interactions, deviat…

cs.CL2025

Towards Conversational AI for Disease Management

Anil Palepu, Valentin Liévin, Wei-Hung Weng +17

While large language models (LLMs) have shown promise in diagnostic dialogue, their capabilities for effective management reasoning - including disease progression, therapeutic res…