76 citations · 78 across the 4 of their papers we have counts for
3 papers · 1 filter
HealthBench Professional: Evaluating Large Language Models on Real Clinician Chats
Rebecca Soskin Hicks, Mikhail Trofimov, Dominick Lim +13
Millions of clinicians use ChatGPT to support clinical care, but evaluations of the most common use cases in model-clinician conversations are limited. We introduce HealthBench Pro…
Advancing Conversational Diagnostic AI with Multimodal Reasoning
Khaled Saab, Jan Freyberg, Chunjong Park +33
Large Language Models (LLMs) have demonstrated great potential for conducting diagnostic conversations but evaluation has been largely limited to language-only interactions, deviat…
Towards Conversational AI for Disease Management
Anil Palepu, Valentin Liévin, Wei-Hung Weng +17
While large language models (LLMs) have shown promise in diagnostic dialogue, their capabilities for effective management reasoning - including disease progression, therapeutic res…