works on

From the 1 of 8 linked papers with an AI index.

most citedFirst, do NOHARM: a medical safety benchmark and randomized study of physician and AI teaming on clinical consultations

6 citations · 8 across the 7 of their papers we have counts for

collaborators

8 papers

cs.CY20266 cited

First, do NOHARM: a medical safety benchmark and randomized study of physician and AI teaming on clinical consultations

David Wu, Fateme Nateghi Haredasht, Saloni Kumar Maharaj +54

The paper introduces NOHARM, a benchmark of 1,100 primary‑care to specialist consultation cases, to evaluate how often large language models and retrieval‑augmented clinical AI too…

cs.AI2026

A global log for medical AI

Ayush Noori, Aaron E. Boussina, Hai Ho Bich +48

Modern computer systems rely on syslog, a universal protocol that records critical events across heterogeneous infrastructure. Medicine's rapidly growing AI stack has no equivalent…

cs.AI2026

Deployment-Centered Evaluation: Predicting Query-Level Rejection Risk in a Clinical LLM System

Alyssa Unell, Miguel Fuentes, Brenna Li +4

Large language models (LLMs) are increasingly integrated into clinical systems, making it essential to evaluate the real-world utility of these systems. However, static benchmarks…

cs.CL2026

CARE: A Conformal Safety Layer for Medical Summarization

Suhana Bedi, Bridget Lin, Anson Y. Zhou +5

Large language models (LLMs) are increasingly used for medical summarization, but their outputs can omit medically important information and introduce unsupported claims. Existing…

cs.CL2026

Quantifying and Mitigating Premature Closure in Frontier LLMs

Rebecca Handler, Suhana Bedi, Nigam Shah

Premature closure, or committing to a conclusion before sufficient information is available, is a recognized contributor to diagnostic error but remains underexamined in large lang…

cs.CY20262 cited

Adoption and Use of LLMs at an Academic Medical Center

Nigam H. Shah, Nerissa Ambers, Abby Pandya +55

While large language models (LLMs) can support clinical documentation needs, standalone tools struggle with "workflow friction" from manual data entry. We developed ChatEHR, a syst…