Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Scaling Clinical Judgment to Evaluate Medical AI
Thomas A. Buckley, Zahir Kanjee, Peter G. Brodeur +14
Blinded physician evaluation has been considered by many to be the gold standard for assessing clinical reasoning in large language models (LLMs). This is difficult to scale; thus,…
cs.AI2025
Teaching large language models to reason like expert diagnosticians
Thomas A. Buckley, Riccardo Conci, Peter G. Brodeur +23
Differential diagnosis is an iterative process that integrates patient information with broader medical knowledge. Clinical case series such as the NEJM Clinicopathologic Conferenc…