8 papers
ConRub-Med: Reinforcement Learning with Consensus Rubrics for Open-Ended Medical Question Answering
Taojie Zhu, Yuan Xia, Tao Sun +8
Reinforcement learning with verifiable rewards has been especially effective in mathematics and coding, where answers can be checked automatically. Many open-ended medical question…
Evolving Interactive Diagnostic Agents in a Virtual Clinical Environment
Pengcheng Qiu, Chaoyi Wu, Junwei Liu +11
We present a framework for training large language models (LLMs) as diagnostic agents with reinforcement learning, enabling them to manage multi-turn interactive diagnostic process…
MLB: A Scenario-Driven Benchmark for Evaluating Large Language Models in Clinical Applications
Qing He, Dongsheng Bi, Jianrong Lu +20
The proliferation of Large Language Models (LLMs) presents transformative potential for healthcare, yet practical deployment is hindered by the absence of frameworks that assess re…
MedDialogRubrics: A Comprehensive Benchmark and Evaluation Framework for Multi-turn Medical Consultations in Large Language Models
Lecheng Gong, Weimin Fang, Ting Yang +9
Medical conversational AI (AI) plays a pivotal role in the development of safer and more effective medical dialogue systems. However, existing benchmarks and evaluation frameworks…
GAPS: A Clinically Grounded, Automated Benchmark for Evaluating AI Clinicians
Xiuyuan Chen, Tao Sun, Dexin Su +37
Current benchmarks for AI clinician systems, often based on multiple-choice exams or manual rubrics, fail to capture the depth, robustness, and safety required for real-world clini…
EHR-R1: A Reasoning-Enhanced Foundational Language Model for Electronic Health Record Analysis
Yusheng Liao, Chaoyi Wu, Junwei Liu +12
Electronic Health Records (EHRs) contain rich yet complex information, and their automated analysis is critical for clinical decision-making. Despite recent advances of large langu…