8 papers
HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents
Qianchu Liu, Sheng Zhang, Guanghui Qin +16
As AI agents become increasingly capable of complex, long-horizon reasoning, rigorous and holistic evaluation is essential for measuring progress toward real-world healthcare appli…
Development, Evaluation, and Deployment of a Multi-Agent System for Thoracic Tumor Board
Tim Ellis-Caleo, Timothy Keyes, Nerissa Ambers +5
Tumor boards are multidisciplinary conferences dedicated to producing actionable patient care recommendations with live review of primary radiology and pathology data. Succinct pat…
RADAR: A Multimodal Benchmark for 3D Image-Based Radiology Report Review
Zhaoyi Sun, Minal Jagtiani, Wen-wai Yim +4
Radiology reports for the same patient examination may contain clinically meaningful discrepancies arising from interpretation differences, reporting variability, or evolving asses…
CoRe-BT: A Multimodal Radiology-Pathology-Text Benchmark for Robust Brain Tumor Typing
Juampablo E. Heras Rivera, Daniel K. Low, Xavier Xiong +5
Accurate brain tumor typing requires integrating heterogeneous clinical evidence, including magnetic resonance imaging (MRI), histopathology, and pathology reports, which are often…
Scaling medical imaging report generation with multimodal reinforcement learning
Qianchu Liu, Sheng Zhang, Guanghui Qin +11
Frontier models have demonstrated remarkable capabilities in understanding and reasoning with natural-language text, but they still exhibit major competency gaps in multimodal unde…
DermaVQA-DAS: Dermatology Assessment Schema (DAS) & Datasets for Closed-Ended Question Answering & Segmentation in Patient-Generated Dermatology Images
Wen-wai Yim, Yujuan Fu, Asma Ben Abacha +4
Recent advances in dermatological image analysis have been driven by large-scale annotated datasets; however, most existing benchmarks focus on dermatoscopic images and lack patien…