5 papers
EHRBench: An Automated and Reliable EHR-based Benchmark for Clinical Decision Making with LLMs
Yuzhang Xie, Keqi Han, Yunpeng Xiao +7
Clinical decision-making (CDM) is central to real-world clinical workflows, where clinicians infer diagnoses, select treatments, or anticipate future health outcomes under incomple…
EpiQAL: Benchmarking Large Language Models in Epidemiological Question Answering and Reasoning
Mingyang Wei, Dehai Min, Zewen Liu +8
Reliable epidemiological reasoning requires synthesizing study evidence to infer disease burden, transmission dynamics, and intervention effects at the population level. Existing m…
Dialogue to Question Generation for Evidence-based Medical Guideline Agent Development
Zongliang Ji, Ziyang Zhang, Xincheng Tan +5
Evidence-based medicine (EBM) is central to high-quality care, but remains difficult to implement in fast-paced primary care settings. Physicians face short consultations, increasi…
Measuring Spiritual Values and Bias of Large Language Models
Songyuan Liu, Ziyang Zhang, Runze Yan +3
Large language models (LLMs) have become integral tool for users from various backgrounds. LLMs, trained on vast corpora, reflect the linguistic and cultural nuances embedded in th…
KERAP: A Knowledge-Enhanced Reasoning Approach for Accurate Zero-shot Diagnosis Prediction Using Multi-agent LLMs
Yuzhang Xie, Hejie Cui, Ziyang Zhang +5
Medical diagnosis prediction plays a critical role in disease detection and personalized healthcare. While machine learning (ML) models have been widely adopted for this task, thei…