Publications (4)
Monitoring Deployed AI Systems in Health Care
Timothy Keyes, Alison Callahan, Abby S. Pandya +18
Post-deployment monitoring of artificial intelligence (AI) systems in health care is essential to ensure their safety, quality, and sustained benefit-and to support governance deci…
MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Suhana Bedi, Hejie Cui, Miguel Fuentes +78
While large language models (LLMs) achieve near-perfect scores on medical licensing exams, these evaluations inadequately reflect the complexity and diversity of real-world clinica…
Zero-Shot Clinical Trial Patient Matching with LLMs
Michael Wornow, Alejandro Lozano, Dev Dash +3
Matching patients to clinical trials is a key unsolved challenge in bringing new drugs to market. Today, identifying patients who meet a trial's eligibility criteria is highly manu…
VeriFact: Verifying Facts in LLM-Generated Clinical Text with Electronic Health Records
Philip Chung, Akshay Swaminathan, Alex J. Goodell +26
Methods to ensure factual accuracy of text generated by large language models (LLM) in clinical medicine are lacking. VeriFact is an artificial intelligence system that combines re…