4 papers
HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents
Qianchu Liu, Sheng Zhang, Guanghui Qin +16
As AI agents become increasingly capable of complex, long-horizon reasoning, rigorous and holistic evaluation is essential for measuring progress toward real-world healthcare appli…
CodeClinic: Evaluating Automation of Coding Skills for Clinical Reasoning Agents
Timothy Ossowski, Xinchi Liu, Danyal Maqbool +6
Clinical reasoning agents based on large language models (LLMs) aim to automate tasks such as intensive care unit (ICU) monitoring and patient state tracking from electronic health…
The Illusion of Readiness in Health AI
Yu Gu, Jingjing Fu, Xiaodong Liu +29
Large language models have demonstrated remarkable performance in a wide range of medical benchmarks. Yet underneath the seemingly promising results lie salient growth areas, espec…
Generative Medical Event Models Improve with Scale
Shane Waxler, Paul Blazek, Davis White +16
Realizing personalized medicine at scale calls for methods that distill insights from longitudinal patient journeys, which can be viewed as a sequence of medical events. Foundation…