4 papers
HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents
Qianchu Liu, Sheng Zhang, Guanghui Qin +16
As AI agents become increasingly capable of complex, long-horizon reasoning, rigorous and holistic evaluation is essential for measuring progress toward real-world healthcare appli…
One Loss to Rule Them All: Marked Time-to-Event for Structured EHR Foundation Models
Zilin Jing, Vincent Jeanselme, Yuta Kobayashi +6
Clinical events captured in Electronic Health Records (EHR) are irregularly sampled and may consist of a mixture of discrete events and numerical measurements, such as laboratory v…
Learning-To-Measure: In-Context Active Feature Acquisition
Yuta Kobayashi, Zilin Jing, Jiayu Yao +2
Active feature acquisition (AFA) is a sequential decision-making problem where the goal is to improve model performance for test instances by adaptively selecting which features to…
FoMoH: A clinically meaningful foundation model evaluation for structured electronic health records
Chao Pang, Vincent Jeanselme, Young Sang Choi +9
Foundation models (FMs) promise to address core limitations of traditional supervised machine learning: (i) reliance on large amounts of labeled data, (ii) task specificity, and (i…