10 papers
The Illusion of Readiness in Health AI
Yu Gu, Jingjing Fu, Xiaodong Liu +29
Large language models have demonstrated remarkable performance in a wide range of medical benchmarks. Yet underneath the seemingly promising results lie salient growth areas, espec…
CancerGUIDE: Cancer Guideline Understanding via Internal Disagreement Estimation
Alyssa Unell, Noel C. F. Codella, Sam Preston +13
The National Comprehensive Cancer Network (NCCN) provides evidence-based guidelines for cancer treatment. Translating complex patient presentations into guideline-compliant treatme…
Demo: Healthcare Agent Orchestrator (HAO) for Patient Summarization in Molecular Tumor Boards
Matthias Blondeel, Noel Codella, Sam Preston +9
Molecular Tumor Boards (MTBs) are multidisciplinary forums where oncology specialists collaboratively assess complex patient cases to determine optimal treatment strategies. A cent…
From Embeddings to Accuracy: Comparing Foundation Models for Radiographic Classification
Xue Li, Jameson Merkow, Noel C. F. Codella +11
Foundation models provide robust embeddings for diverse tasks, including medical imaging. We evaluate embeddings from seven general and medical-specific foundation models (e.g., De…
Sequential Diagnosis with Language Models
Harsha Nori, Mayank Daswani, Christopher Kelly +12
Artificial intelligence holds great promise for expanding access to expert medical knowledge and reasoning. However, most evaluations of language models rely on static vignettes an…
MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Suhana Bedi, Hejie Cui, Miguel Fuentes +78
While large language models (LLMs) achieve near-perfect scores on medical licensing exams, these evaluations inadequately reflect the complexity and diversity of real-world clinica…