collaborators

9 papers

cs.CL2026

RESPClinBench: Benchmarking Multimodal Clinical Decision-Making and Longitudinal Disease Management in Respiratory Specialty Care

Mouxiao Bian, Zhi Chen, Ruiyao Chen +9

Background: Respiratory specialty care requires multimodal interpretation, longitudinal risk assessment, guideline-concordant intervention, and whole-course management, which are p…

cs.CL2026

CardioBench: A Real-World Data Benchmark for Evaluating Large Language Models in Clinically Authentic Cardiovascular Care Scenarios

Xiao Li, Mouxiao Bian, Zhaodi Wu +8

Background: Most medical large language model (LLM) benchmarks focus on examination knowledge or isolated tasks and may not reflect the longitudinal, multimodal, and safety-critica…

cs.AI2026

MedRealMM: A Real-World Multimodal Benchmark for Chinese Online Medical Consultation

Runhan Shi, Quan Zhou, Yuqian Xu +14

The paper presents MedRealMM, a large-scale benchmark of real Chinese online medical consultations that includes both text and patient-uploaded images, and evaluates how well large…

cs.CL2026

MedBench v5: A Dynamic, Process-Oriented, and Hallucination-Aware Benchmark for Clinical Multimodal Models

Jinru Ding, Chuchu Jiang, Lu Lu +12

Existing medical AI benchmarks lack process visibility, atomic skill evaluation, and integrated hallucination detection. We introduce MedBench v5, a redesigned benchmark for clinic…

cs.CL2025

MedBench v4: A Robust and Scalable Benchmark for Evaluating Chinese Medical Language Models, Multimodal Models, and Intelligent Agents

Jinru Ding, Lu Lu, Chao Ding +15

Recent advances in medical large language models (LLMs), multimodal models, and agents demand evaluation frameworks that reflect real clinical workflows and safety constraints. We…

cs.CL2025

Human-Level and Beyond: Benchmarking Large Language Models Against Clinical Pharmacists in Prescription Review

Yan Yang, Mouxiao Bian, Peiling Li +10

The rapid advancement of large language models (LLMs) has accelerated their integration into clinical decision support, particularly in prescription review. To enable systematic an…