4 papers
Scaling Clinical Judgment to Evaluate Medical AI
Thomas A. Buckley, Zahir Kanjee, Peter G. Brodeur +14
Blinded physician evaluation has been considered by many to be the gold standard for assessing clinical reasoning in large language models (LLMs). This is difficult to scale; thus,…
Teaching large language models to reason like expert diagnosticians
Thomas A. Buckley, Riccardo Conci, Peter G. Brodeur +23
Differential diagnosis is an iterative process that integrates patient information with broader medical knowledge. Clinical case series such as the NEJM Clinicopathologic Conferenc…
Superhuman performance of a large language model on the reasoning tasks of a physician
Peter G. Brodeur, Thomas A. Buckley, Zahir Kanjee +22
A seminal paper published by Ledley and Lusted in 1959 introduced complex clinical diagnostic reasoning cases as the gold standard for the evaluation of expert medical computing sy…
Performance of large language models in numerical vs. semantic medical knowledge: Benchmarking on evidence-based Q&As
Eden Avnat, Michal Levy, Daniel Herstain +12
Clinical problem-solving requires processing of semantic medical knowledge such as illness scripts and numerical medical knowledge of diagnostic tests for evidence-based decision-m…