4 papers · 1 filter
CliniCARE-Bench: Clinical Calibrated Audit of Medical Reasoning in EHR
Veronica Chatrath, Bryan Zhu, George Pu +16
Large language models perform strongly on medical knowledge benchmarks, but reliable clinical deployment requires agents to conduct defensible investigations over heterogeneous, lo…
PSEBench: A Controllable and Verifiable Benchmark for Evaluating LLMs in Patient Safety Event Triage
Keqi Han, Ryan Young, Annabel Strauss +7
Patient safety event triage, determining whether a clinical event is reportable under jurisdiction-specific policy, is a high-stakes task typically performed manually by patient sa…
Towards a Virtual Neuroscientist: Autonomous Neuroimaging Analysis via Multi-Agent Collaboration
Keqi Han, Songlin Zhao, Yao Su +4
Transforming neuroimaging data into clinically actionable biomarkers is a knowledge-intensive and labor-intensive process. Standardized workflows such as fMRIPrep have improved rob…
EHRBench: An Automated and Reliable EHR-based Benchmark for Clinical Decision Making with LLMs
Yuzhang Xie, Keqi Han, Yunpeng Xiao +7
Clinical decision-making (CDM) is central to real-world clinical workflows, where clinicians infer diagnoses, select treatments, or anticipate future health outcomes under incomple…