activity
20242026
collaborators

5 papers

cs.AI2026

MedCalc-Pro: Solving Complex Medical Calculations with LLM Agents

Siran Zhao, Ruihui Hou, Ziyue Huai +2

Current benchmarks for evaluating large language models (LLMs) in medical calculation are largely based on simplified settings, where each patient case corresponds to a single calc…

cs.AI2026

ClinicalMC: A Benchmark for Multi-Course Clinical Decision-Making with Large Language Models

Ruihui Hou, Siyi Zhu, Ziyue Huai +4

Large language models (LLMs) have been widely adopted in healthcare, yet they still encounter significant challenges in complex clinical decision-making scenarios. Existing benchma…

cs.AI2026

CAREAgent: Clinical Agent with Structured Reasoning and Tool-Integrated for Order Generation

Ruihui Hou, Ziyue Huai, Chennuo Zhang +5

Clinical order generation serves as a critical bridge between clinical decision-making and real-world practice, translating medical decisions into concrete and executable orders. E…

cs.CL2025

CMQCIC-Bench: A Chinese Benchmark for Evaluating Large Language Models in Medical Quality Control Indicator Calculation

Guangya Yu, Yanhao Li, Zongying Jiang +9

Medical quality control indicators are essential to assess the qualifications of healthcare institutions for medical services. With the impressive performance of large language mod…

cs.AI2024

MSDiagnosis: A Benchmark for Evaluating Large Language Models in Multi-Step Clinical Diagnosis

Ruihui Hou, Shencheng Chen, Yongqi Fan +5

Clinical diagnosis is critical in medical practice, typically requiring a continuous and evolving process that includes primary diagnosis, differential diagnosis, and final diagnos…