5 papers
Scaling Clinical Judgment to Evaluate Medical AI
Thomas A. Buckley, Zahir Kanjee, Peter G. Brodeur +14
Blinded physician evaluation has been considered by many to be the gold standard for assessing clinical reasoning in large language models (LLMs). This is difficult to scale; thus,…
How Seemingly Inconsequential Design Choices Dictate Performance of LLMs in Pathology
Kian R. Weihrauch, Thomas A. Buckley, William Lotter +1
General-purpose large language models (LLMs) are routinely used as baselines when evaluating specialized pathology models on whole-slide images (WSIs). Because WSIs exceed contempo…
Navigating Gigapixel Pathology Images with Large Multimodal Models
Thomas A. Buckley, Kian R. Weihrauch, Katherine Latham +3
Recent advances in large multimodal models have allowed for the development of interactive chat models that can converse and reason about pathology whole-slide images (WSIs). Howev…
Teaching large language models to reason like expert diagnosticians
Thomas A. Buckley, Riccardo Conci, Peter G. Brodeur +23
Differential diagnosis is an iterative process that integrates patient information with broader medical knowledge. Clinical case series such as the NEJM Clinicopathologic Conferenc…
Superhuman performance of a large language model on the reasoning tasks of a physician
Peter G. Brodeur, Thomas A. Buckley, Zahir Kanjee +22
A seminal paper published by Ledley and Lusted in 1959 introduced complex clinical diagnostic reasoning cases as the gold standard for the evaluation of expert medical computing sy…