most citedMedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

10 citations · 10 across the 4 of their papers we have counts for

collaborators

5 papers

cs.AI2026

Development, Evaluation, and Deployment of a Multi-Agent System for Thoracic Tumor Board

Tim Ellis-Caleo, Timothy Keyes, Nerissa Ambers +5

Tumor boards are multidisciplinary conferences dedicated to producing actionable patient care recommendations with live review of primary radiology and pathology data. Succinct pat…

cs.CY2026

Deployment and Evaluation of an EHR-integrated, Large Language Model-Powered Tool to Triage Surgical Patients

Jane Wang, Timothy Keyes, April S Liang +10

Surgical co-management (SCM) is an evidence-based model in which hospitalists jointly manage medically complex perioperative patients alongside surgical teams. Despite its clinical…

cs.CY2026

Adoption and Use of LLMs at an Academic Medical Center

Nigam H. Shah, Nerissa Ambers, Abby Pandya +55

While large language models (LLMs) can support clinical documentation needs, standalone tools struggle with "workflow friction" from manual data entry. We developed ChatEHR, a syst…

q-bio.OT2026

Monitoring Deployed AI Systems in Health Care

Timothy Keyes, Alison Callahan, Abby S. Pandya +18

Post-deployment monitoring of artificial intelligence (AI) systems in health care is essential to ensure their safety, quality, and sustained benefit-and to support governance deci…

cs.CL202510 cited

MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

Suhana Bedi, Hejie Cui, Miguel Fuentes +78

While large language models (LLMs) achieve near-perfect scores on medical licensing exams, these evaluations inadequately reflect the complexity and diversity of real-world clinica…