4 papers
MedFactEval and MedAgentBrief: A Framework and Workflow for Generating and Evaluating Factual Clinical Summaries
François Grolleau, Emily Alsentzer, Timothy Keyes +17
Evaluating factual accuracy in Large Language Model (LLM)-generated clinical text is a critical barrier to adoption, as expert review is unscalable for the continuous quality assur…
FactEHR: A Dataset for Evaluating Factuality in Clinical Notes Using LLMs
Monica Munnangi, Akshay Swaminathan, Jason Alan Fries +8
Verifying and attributing factual claims is essential for the safe and effective use of large language models (LLMs) in healthcare. A core component of factuality evaluation is fac…
MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Suhana Bedi, Hejie Cui, Miguel Fuentes +78
While large language models (LLMs) achieve near-perfect scores on medical licensing exams, these evaluations inadequately reflect the complexity and diversity of real-world clinica…
VeriFact: Verifying Facts in LLM-Generated Clinical Text with Electronic Health Records
Philip Chung, Akshay Swaminathan, Alex J. Goodell +26
Methods to ensure factual accuracy of text generated by large language models (LLM) in clinical medicine are lacking. VeriFact is an artificial intelligence system that combines re…