cs.CY2026
First, do NOHARM: a medical safety benchmark and randomized study of physician and AI teaming on clinical consultations
David Wu, Fateme Nateghi Haredasht, Saloni Kumar Maharaj +54
The paper introduces NOHARM, a benchmark of 1,100 primary‑care to specialist consultation cases, to evaluate how often large language models and retrieval‑augmented clinical AI too…
#medical safety#large language models#clinical decision support#human‑ai teaming