9 papers
When Multiple Scripts Matter: Evaluating ASR in Clinical Settings
Jean Seo, Minkyu Kim, Jeonguk Lee +3
Automatic speech recognition (ASR) in non-English clinical settings is challenged by multiscript variability, where the same term may appear in multiple valid orthographic forms. C…
Evaluating the Pre-Consultation Ability of LLMs using Diagnostic Guidelines
Jean Seo, Gibaeg Kim, Kihun Shin +5
We introduce EPAG, a benchmark dataset and framework designed for Evaluating the Pre-consultation Ability of LLMs using diagnostic Guidelines. LLMs are evaluated directly through H…
H-DDx: A Hierarchical Evaluation Framework for Differential Diagnosis
Seungseop Lim, Gibaeg Kim, Hyunkyung Lee +4
An accurate differential diagnosis (DDx) is essential for patient care, shaping therapeutic decisions and influencing outcomes. Recently, Large Language Models (LLMs) have emerged…
Format Inertia: A Failure Mechanism of LLMs in Medical Pre-Consultation
Seungseop Lim, Gibaeg Kim, Wooseok Han +4
Recent advances in Large Language Models (LLMs) have brought significant improvements to various service domains, including chatbots and medical pre-consultation applications. In t…
Taxonomy of Comprehensive Safety for Clinical Agents
Jean Seo, Hyunkyung Lee, Gibaeg Kim +5
Safety is a paramount concern in clinical chatbot applications, where inaccurate or harmful responses can lead to serious consequences. Existing methods--such as guardrails and too…
DeltaSHAP: Explaining Prediction Evolutions in Online Patient Monitoring with Shapley Values
Changhun Kim, Yechan Mun, Sangchul Hahn +1
This study proposes DeltaSHAP, a novel explainable artificial intelligence (XAI) algorithm specifically designed for online patient monitoring systems. In clinical environments, di…