2 citations · 2 across the 3 of their papers we have counts for
1 paper · 1 filter
Thomas A. Buckley, Zahir Kanjee, Peter G. Brodeur +14
Blinded physician evaluation has been considered by many to be the gold standard for assessing clinical reasoning in large language models (LLMs). This is difficult to scale; thus,…