3 papers
cs.AI2026
HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents
Qianchu Liu, Sheng Zhang, Guanghui Qin +16
As AI agents become increasingly capable of complex, long-horizon reasoning, rigorous and holistic evaluation is essential for measuring progress toward real-world healthcare appli…
cs.AI2025
The Illusion of Readiness in Health AI
Yu Gu, Jingjing Fu, Xiaodong Liu +29
Large language models have demonstrated remarkable performance in a wide range of medical benchmarks. Yet underneath the seemingly promising results lie salient growth areas, espec…
cs.LG2025
Generative Medical Event Models Improve with Scale
Shane Waxler, Paul Blazek, Davis White +16
Realizing personalized medicine at scale calls for methods that distill insights from longitudinal patient journeys, which can be viewed as a sequence of medical events. Foundation…