Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
MedCalc-Pro: Solving Complex Medical Calculations with LLM Agents
Siran Zhao, Ruihui Hou, Ziyue Huai +2
Current benchmarks for evaluating large language models (LLMs) in medical calculation are largely based on simplified settings, where each patient case corresponds to a single calc…
cs.AI2026
ClinicalMC: A Benchmark for Multi-Course Clinical Decision-Making with Large Language Models
Ruihui Hou, Siyi Zhu, Ziyue Huai +4
Large language models (LLMs) have been widely adopted in healthcare, yet they still encounter significant challenges in complex clinical decision-making scenarios. Existing benchma…