most citedMedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

10 citations · 10 across the 4 of their papers we have counts for

collaborators

5 papers

cs.CL2026

Large Language Models for Large-Scale, Rigorous Qualitative Analysis in Applied Health Services Research

Sasha Ronaghi, Emma-Louise Aveling, Maria Levis +3

Large language models (LLMs) show promise for improving the efficiency of qualitative analysis in large, multi-site health-services research. Yet methodological guidance for LLM in…

cs.CL2025

Retrieval-Augmented Guardrails for AI-Drafted Patient-Portal Messages: Error Taxonomy Construction and Large-Scale Evaluation

Wenyuan Chen, Fateme Nateghi Haredasht, Kameron C. Black +4

Asynchronous patient-clinician messaging via EHR portals is a growing source of clinician workload, prompting interest in large language models (LLMs) to assist with draft response…

cs.CL2025

MedFactEval and MedAgentBrief: A Framework and Workflow for Generating and Evaluating Factual Clinical Summaries

François Grolleau, Emily Alsentzer, Timothy Keyes +17

Evaluating factual accuracy in Large Language Model (LLM)-generated clinical text is a critical barrier to adoption, as expert review is unscalable for the continuous quality assur…

cs.CL202510 cited

MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

Suhana Bedi, Hejie Cui, Miguel Fuentes +78

While large language models (LLMs) achieve near-perfect scores on medical licensing exams, these evaluations inadequately reflect the complexity and diversity of real-world clinica…

cs.AI2025

TIMER: Temporal Instruction Modeling and Evaluation for Longitudinal Clinical Records

Hejie Cui, Alyssa Unell, Bowen Chen +4

Large language models (LLMs) have emerged as promising tools for assisting in medical tasks, yet processing Electronic Health Records (EHRs) presents unique challenges due to their…