1 citations · 1 across the 1 of their papers we have counts for
2 papers
cs.AI2026
Scaling Clinical Judgment to Evaluate Medical AI
Thomas A. Buckley, Zahir Kanjee, Peter G. Brodeur +14
Blinded physician evaluation has been considered by many to be the gold standard for assessing clinical reasoning in large language models (LLMs). This is difficult to scale; thus,…
cs.AI2025★ 1 cited
Teaching large language models to reason like expert diagnosticians
Thomas A. Buckley, Riccardo Conci, Peter G. Brodeur +23
Differential diagnosis is an iterative process that integrates patient information with broader medical knowledge. Clinical case series such as the NEJM Clinicopathologic Conferenc…