10 citations · 18 across the 5 of their papers we have counts for
5 papers
Measurement to Meaning: A Validity-Centered Framework for AI Evaluation
Olawale Salaudeen, Anka Reuel, Ahmed Ahmed +6
While the capabilities and utility of AI systems have advanced, rigorous norms for evaluating these systems have lagged. Grand claims, such as models achieving general reasoning ca…
Disentangling Reasoning and Knowledge in Medical Large Language Models
Rahul Thapa, Qingyang Wu, Kevin Wu +11
Medical reasoning in large language models (LLMs) aims to emulate clinicians' diagnostic thinking, but current benchmarks such as MedQA-USMLE, MedMCQA, and PubMedQA often mix reaso…
The Optimization Paradox in Clinical AI Multi-Agent Systems
Suhana Bedi, Iddah Mlauzi, Daniel Shin +2
Multi-agent artificial intelligence systems are increasingly deployed in clinical settings, yet the relationship between component-level optimization and system-wide performance re…
MedHELM: Holistic Evaluation of Large Language Models for Medical Tasks
Suhana Bedi, Hejie Cui, Miguel Fuentes +78
While large language models (LLMs) achieve near-perfect scores on medical licensing exams, these evaluations inadequately reflect the complexity and diversity of real-world clinica…
Distilling Large Language Models for Efficient Clinical Information Extraction
Karthik S. Vedula, Annika Gupta, Akshay Swaminathan +3
Large language models (LLMs) excel at clinical information extraction but their computational demands limit practical deployment. Knowledge distillation--the process of transferrin…