4 papers
Disentangling Reasoning and Knowledge in Medical Large Language Models
Rahul Thapa, Qingyang Wu, Kevin Wu +11
Medical reasoning in large language models (LLMs) aims to emulate clinicians' diagnostic thinking, but current benchmarks such as MedQA-USMLE, MedMCQA, and PubMedQA often mix reaso…
MedCaseReasoning: Evaluating and learning diagnostic reasoning from clinical case reports
Kevin Wu, Eric Wu, Rahul Thapa +7
Doctors and patients alike increasingly use Large Language Models (LLMs) to diagnose clinical cases. However, unlike domains such as math or coding, where correctness can be object…
AutoRedTeamer: Autonomous Red Teaming with Lifelong Attack Integration
Andy Zhou, Kevin Wu, Francesco Pinto +7
As large language models (LLMs) become increasingly capable, security and safety evaluation are crucial. While current red teaming approaches have made strides in assessing LLM vul…
ClashEval: Quantifying the tug-of-war between an LLM's internal prior and external evidence
Kevin Wu, Eric Wu, James Zou
Retrieval augmented generation (RAG) is frequently used to mitigate hallucinations and provide up-to-date knowledge for large language models (LLMs). However, given that document r…