activity
20242026
most citedHow well do LLMs cite relevant medical references? An evaluation framework and analyses

41 citations · 50 across the 8 of their papers we have counts for

collaborators
Showing cs.CLShow all

6 papers · 1 filter

cs.CL2026

MedArena: Comparing LLMs for Medicine-in-the-Wild Clinician Preferences

Eric Wu, Kevin Wu, Jason Hom +11

Large language models (LLMs) are increasingly central to clinician workflows, spanning clinical decision support, medical education, and patient communication. However, current eva…

cs.CL20251 cited

Disentangling Reasoning and Knowledge in Medical Large Language Models

Rahul Thapa, Qingyang Wu, Kevin Wu +11

Medical reasoning in large language models (LLMs) aims to emulate clinicians' diagnostic thinking, but current benchmarks such as MedQA-USMLE, MedMCQA, and PubMedQA often mix reaso…

cs.CL20255 cited

MedCaseReasoning: Evaluating and learning diagnostic reasoning from clinical case reports

Kevin Wu, Eric Wu, Rahul Thapa +7

Doctors and patients alike increasingly use Large Language Models (LLMs) to diagnose clinical cases. However, unlike domains such as math or coding, where correctness can be object…

cs.CL20242 cited

FineTuneBench: How well do commercial fine-tuning APIs infuse knowledge into LLMs?

Eric Wu, Kevin Wu, James Zou

There is great interest in fine-tuning frontier large language models (LLMs) to inject new information and update existing knowledge. While commercial LLM fine-tuning APIs from pro…

cs.CL2024

ClashEval: Quantifying the tug-of-war between an LLM's internal prior and external evidence

Kevin Wu, Eric Wu, James Zou

Retrieval augmented generation (RAG) is frequently used to mitigate hallucinations and provide up-to-date knowledge for large language models (LLMs). However, given that document r…

cs.CL202441 cited

How well do LLMs cite relevant medical references? An evaluation framework and analyses

Kevin Wu, Eric Wu, Ally Cassasola +7

Large language models (LLMs) are currently being used to answer medical questions across a variety of clinical domains. Recent top-performing commercial LLMs, in particular, are al…