papers

Publications (14)

cs.HC2026

A prospective clinical feasibility study of a conversational diagnostic AI in an ambulatory primary care clinic

Peter Brodeur, Jacob M. Koshy, Anil Palepu +45

Large language model (LLM)-based AI systems have shown promise for patient-facing diagnostic and management conversations in simulated settings. Translating these systems into clin…

cs.AI2026

A global log for medical AI

Ayush Noori, Aaron E. Boussina, Hai Ho Bich +48

Modern computer systems rely on syslog, a universal protocol that records critical events across heterogeneous infrastructure. Medicine's rapidly growing AI stack has no equivalent…

cs.CL2026

ER-Reason: A Benchmark Dataset for LLM Clinical Reasoning in the Emergency Room

Nikita Mehandru, Niloufar Golchini, Namrata Garg +9

Existing benchmarks for evaluating the clinical reasoning capabilities of large language models (LLMs) often lack a clear definition of "clinical reasoning" as a construct, fail to…

cs.AI2025

One Patient, Many Contexts: Scaling Medical AI with Contextual Intelligence

Michelle M. Li, Ben Y. Reis, Adam Rodman +6

Medical AI, including clinical language models, vision-language models, and multimodal health record models, already summarizes notes, answers questions, and supports decisions. Th…

cs.CV2024

Multimodal Foundation Models Exploit Text to Make Medical Image Predictions

Thomas Buckley, James A. Diao, Pranav Rajpurkar +2

Multimodal foundation models have shown compelling but conflicting performance in medical image interpretation. However, the mechanisms by which these models integrate and prioriti…

cs.CY2026

First, do NOHARM: a medical safety benchmark and randomized study of physician and AI teaming on clinical consultations

David Wu, Fateme Nateghi Haredasht, Saloni Kumar Maharaj +54

The paper introduces NOHARM, a benchmark of 1,100 primary‑care to specialist consultation cases, to evaluate how often large language models and retrieval‑augmented clinical AI too…

#medical safety#large language models#clinical decision support#human‑ai teaming
cs.AI2026

Teaching large language models to reason like expert diagnosticians

Thomas A. Buckley, Riccardo Conci, Peter G. Brodeur +23

Differential diagnosis is an iterative process that integrates patient information with broader medical knowledge. Clinical case series such as the NEJM Clinicopathologic Conferenc…

cs.CL2025

Towards Conversational AI for Disease Management

Anil Palepu, Valentin Liévin, Wei-Hung Weng +17

While large language models (LLMs) have shown promise in diagnostic dialogue, their capabilities for effective management reasoning - including disease progression, therapeutic res…

cs.AI2025

Towards physician-centered oversight of conversational diagnostic AI

Elahe Vedadi, David Barrett, Natalie Harris +32

Recent work has demonstrated the promise of conversational AI systems for diagnostic dialogue. However, real-world assurance of patient safety means that providing individual diagn…

cs.CV2026

RadGame: An AI-Powered Platform for Radiology Education

Mohammed Baharoon, Siavash Raissi, John S. Jun +29

We introduce RadGame, an AI-powered gamified platform for radiology education that targets two core skills: localizing findings and generating reports. Traditional radiology traini…

cs.AI2026

Towards Conversational Medical AI with Eyes, Ears and a Voice

Meet Shah, Jason Gusdorf, Anil Palepu +50

The practice of medicine relies not only upon skillful dialogue but also on the nuanced exchange and interpretation of rich auditory and visual cues between doctors and patients. B…

cs.CL2026

BRIDGE: Benchmarking Large Language Models for Understanding Real-world Clinical Practice Text

Jiageng Wu, Bowen Gu, Ren Zhou +14

Large language models (LLMs) hold great promise for medical applications and are evolving rapidly, with new models being released at an accelerated pace. However, benchmarking on l…

cs.AI2025

Superhuman performance of a large language model on the reasoning tasks of a physician

Peter G. Brodeur, Thomas A. Buckley, Zahir Kanjee +22

A seminal paper published by Ledley and Lusted in 1959 introduced complex clinical diagnostic reasoning cases as the gold standard for the evaluation of expert medical computing sy…

cs.CL2025

Advancing Conversational Diagnostic AI with Multimodal Reasoning

Khaled Saab, Jan Freyberg, Chunjong Park +33

Large Language Models (LLMs) have demonstrated great potential for conducting diagnostic conversations but evaluation has been largely limited to language-only interactions, deviat…