benchmark 1clinical decision-making 1electronic health records 1large language models 1longitudinal data 1temporal reasoning 1
From the 1 of 3 linked papers with an AI index.
3 papers
cs.AI2026
LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making
Zihan Xu, Yanzhen Chen, Xiaocheng Zhang +4
The paper presents LongMedBench, a benchmark built from MIMIC-IV electronic health records that evaluates medical agents on long-horizon clinical decision-making across multiple vi…
cs.CV2025
Hulu-Med: A Transparent Generalist Model towards Holistic Medical Vision-Language Understanding
Songtao Jiang, Yuan Wang, Sibo Song +22
Real-world clinical decision-making requires integrating heterogeneous data, including medical text, 2D images, 3D volumes, and videos, while existing AI systems fail to unify all…
cs.CL2025
Med-U1: Incentivizing Unified Medical Reasoning in LLMs via Large-scale Reinforcement Learning
Xiaotian Zhang, Yuan Wang, Zhaopeng Feng +6
Medical Question-Answering (QA) encompasses a broad spectrum of tasks, including multiple choice questions (MCQ), open-ended text generation, and complex computational reasoning. D…