5 citations · 8 across the 3 of their papers we have counts for
4 papers · 1 filter
MedArena: Comparing LLMs for Medicine-in-the-Wild Clinician Preferences
Eric Wu, Kevin Wu, Jason Hom +11
Large language models (LLMs) are increasingly central to clinician workflows, spanning clinical decision support, medical education, and patient communication. However, current eva…
Disentangling Reasoning and Knowledge in Medical Large Language Models
Rahul Thapa, Qingyang Wu, Kevin Wu +11
Medical reasoning in large language models (LLMs) aims to emulate clinicians' diagnostic thinking, but current benchmarks such as MedQA-USMLE, MedMCQA, and PubMedQA often mix reaso…
MedCaseReasoning: Evaluating and learning diagnostic reasoning from clinical case reports
Kevin Wu, Eric Wu, Rahul Thapa +7
Doctors and patients alike increasingly use Large Language Models (LLMs) to diagnose clinical cases. However, unlike domains such as math or coding, where correctness can be object…
FineTuneBench: How well do commercial fine-tuning APIs infuse knowledge into LLMs?
Eric Wu, Kevin Wu, James Zou
There is great interest in fine-tuning frontier large language models (LLMs) to inject new information and update existing knowledge. While commercial LLM fine-tuning APIs from pro…