7 papers
HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents
Qianchu Liu, Sheng Zhang, Guanghui Qin +16
As AI agents become increasingly capable of complex, long-horizon reasoning, rigorous and holistic evaluation is essential for measuring progress toward real-world healthcare appli…
CodeClinic: Evaluating Automation of Coding Skills for Clinical Reasoning Agents
Timothy Ossowski, Xinchi Liu, Danyal Maqbool +6
Clinical reasoning agents based on large language models (LLMs) aim to automate tasks such as intensive care unit (ICU) monitoring and patient state tracking from electronic health…
Scaling medical imaging report generation with multimodal reinforcement learning
Qianchu Liu, Sheng Zhang, Guanghui Qin +11
Frontier models have demonstrated remarkable capabilities in understanding and reasoning with natural-language text, but they still exhibit major competency gaps in multimodal unde…
RiskCueBench: Benchmarking Anticipatory Reasoning from Early Risk Cues in Video-Language Models
Sha Luo, Yogesh Prabhu, Timothy Ossowski +2
With the rapid growth of video centered social media, the ability to anticipate risky events from visual data is a promising direction for ensuring public safety and preventing rea…
COMMA: A Communicative Multimodal Multi-Agent Benchmark
Timothy Ossowski, Danyal Maqbool, Jixuan Chen +3
The rapid advances of multimodal agents built on large foundation models have largely overlooked their potential for language-based communication between agents in collaborative ta…
OctoMed: Data Recipes for State-of-the-Art Multimodal Medical Reasoning
Timothy Ossowski, Sheng Zhang, Qianchu Liu +5
High-quality and carefully curated data is a cornerstone of training medical large language models, as it directly impacts both generalization and robustness to unseen clinical tas…