1 citations · 1 across the 1 of their papers we have counts for
1 paper
Max Lamparth, Declan Grabb, Amy Franks +8
Current medical language model (LM) benchmarks often over-simplify the complexities of day-to-day clinical practice tasks and instead rely on evaluating LMs on multiple-choice boar…