most citedMedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

10 citations · 18 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CY20254 cited

Measurement to Meaning: A Validity-Centered Framework for AI Evaluation

Olawale Salaudeen, Anka Reuel, Ahmed Ahmed +6

While the capabilities and utility of AI systems have advanced, rigorous norms for evaluating these systems have lagged. Grand claims, such as models achieving general reasoning ca…

cs.CL20251 cited

Disentangling Reasoning and Knowledge in Medical Large Language Models

Rahul Thapa, Qingyang Wu, Kevin Wu +11

Medical reasoning in large language models (LLMs) aims to emulate clinicians' diagnostic thinking, but current benchmarks such as MedQA-USMLE, MedMCQA, and PubMedQA often mix reaso…

cs.AI20251 cited

The Optimization Paradox in Clinical AI Multi-Agent Systems

Suhana Bedi, Iddah Mlauzi, Daniel Shin +2

Multi-agent artificial intelligence systems are increasingly deployed in clinical settings, yet the relationship between component-level optimization and system-wide performance re…

cs.CL202510 cited

MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks

Suhana Bedi, Hejie Cui, Miguel Fuentes +78

While large language models (LLMs) achieve near-perfect scores on medical licensing exams, these evaluations inadequately reflect the complexity and diversity of real-world clinica…

cs.CL20242 cited

Distilling Large Language Models for Efficient Clinical Information Extraction

Karthik S. Vedula, Annika Gupta, Akshay Swaminathan +3

Large language models (LLMs) excel at clinical information extraction but their computational demands limit practical deployment. Knowledge distillation--the process of transferrin…