5 papers
PSEBench: A Controllable and Verifiable Benchmark for Evaluating LLMs in Patient Safety Event Triage
Keqi Han, Ryan Young, Annabel Strauss +7
Patient safety event triage, determining whether a clinical event is reportable under jurisdiction-specific policy, is a high-stakes task typically performed manually by patient sa…
EHRBench: An Automated and Reliable EHR-based Benchmark for Clinical Decision Making with LLMs
Yuzhang Xie, Keqi Han, Yunpeng Xiao +7
Clinical decision-making (CDM) is central to real-world clinical workflows, where clinicians infer diagnoses, select treatments, or anticipate future health outcomes under incomple…
Beyond MedQA: Towards Real-world Clinical Decision Making in the Era of LLMs
Yunpeng Xiao, Carl Yang, Mark Mai +2
Large language models (LLMs) show promise for clinical use. They are often evaluated using datasets such as MedQA. However, Many medical datasets, such as MedQA, rely on simplified…
KERAP: A Knowledge-Enhanced Reasoning Approach for Accurate Zero-shot Diagnosis Prediction Using Multi-agent LLMs
Yuzhang Xie, Hejie Cui, Ziyang Zhang +5
Medical diagnosis prediction plays a critical role in disease detection and personalized healthcare. While machine learning (ML) models have been widely adopted for this task, thei…
Piecing It All Together: Verifying Multi-Hop Multimodal Claims
Haoran Wang, Aman Rangapur, Xiongxiao Xu +4
Existing claim verification datasets often do not require systems to perform complex reasoning or effectively interpret multimodal evidence. To address this, we introduce a new tas…