From the 1 of 2 linked papers with an AI index.
6 citations · 6 across the 1 of their papers we have counts for
2 papers
cs.CY2026★ 6 cited
First, do NOHARM: a medical safety benchmark and randomized study of physician and AI teaming on clinical consultations
David Wu, Fateme Nateghi Haredasht, Saloni Kumar Maharaj +54
The paper introduces NOHARM, a benchmark of 1,100 primary‑care to specialist consultation cases, to evaluate how often large language models and retrieval‑augmented clinical AI too…
cs.CL2025
FactEHR: A Dataset for Evaluating Factuality in Clinical Notes Using LLMs
Monica Munnangi, Akshay Swaminathan, Jason Alan Fries +8
Verifying and attributing factual claims is essential for the safe and effective use of large language models (LLMs) in healthcare. A core component of factuality evaluation is fac…