collaborators
Showing cs.AIShow all

5 papers · 1 filter

cs.AI2026

HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents

Qianchu Liu, Sheng Zhang, Guanghui Qin +16

As AI agents become increasingly capable of complex, long-horizon reasoning, rigorous and holistic evaluation is essential for measuring progress toward real-world healthcare appli…

cs.AI2026

CodeClinic: Evaluating Automation of Coding Skills for Clinical Reasoning Agents

Timothy Ossowski, Xinchi Liu, Danyal Maqbool +6

Clinical reasoning agents based on large language models (LLMs) aim to automate tasks such as intensive care unit (ICU) monitoring and patient state tracking from electronic health…

cs.AI2025

COMMA: A Communicative Multimodal Multi-Agent Benchmark

Timothy Ossowski, Danyal Maqbool, Jixuan Chen +3

The rapid advances of multimodal agents built on large foundation models have largely overlooked their potential for language-based communication between agents in collaborative ta…

cs.AI2025

OctoMed: Data Recipes for State-of-the-Art Multimodal Medical Reasoning

Timothy Ossowski, Sheng Zhang, Qianchu Liu +5

High-quality and carefully curated data is a cornerstone of training medical large language models, as it directly impacts both generalization and robustness to unseen clinical tas…

cs.AI2025

X-Reasoner: Towards Generalizable Reasoning Across Modalities and Domains

Qianchu Liu, Sheng Zhang, Guanghui Qin +9

Recent proprietary models (e.g., o3) have begun to demonstrate strong multimodal reasoning capabilities. Yet, most existing open-source research concentrates on training text-only…