large language models 3benchmark 1clinical decision-making 1data synthesis 1electronic health records 1humanities 1humanities and social sciences 1instruction tuning 1longitudinal data 1preference alignment 1quality evaluation 1social sciences 1
From the 3 of 11 linked papers with an AI index.
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making
Zihan Xu, Yanzhen Chen, Xiaocheng Zhang +4
The paper presents LongMedBench, a benchmark built from MIMIC-IV electronic health records that evaluates medical agents on long-horizon clinical decision-making across multiple vi…
cs.AI2025
Reinforcement Learning with Rubric Anchors
Zenan Huang, Yihong Zhuang, Guoshan Lu +18
Reinforcement Learning from Verifiable Rewards (RLVR) has emerged as a powerful paradigm for enhancing Large Language Models (LLMs), exemplified by the success of OpenAI's o-series…