2 papers
cs.AI2026
MedMCP-Calc: Benchmarking LLMs for Realistic Medical Calculator Scenarios via MCP Integration
Yakun Zhu, Yutong Huang, Shengqian Qin +3
Medical calculators are fundamental to quantitative, evidence-based clinical practice. However, their real-world use is an adaptive, multi-stage process, requiring proactive EHR da…
cs.CL2025
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models
Yakun Zhu, Zhongzhen Huang, Linjie Mu +6
The emergence of groundbreaking large language models capable of performing complex reasoning tasks holds significant promise for addressing various scientific challenges, includin…