3 citations · 10 across the 24 of their papers we have counts for
1 paper · 2 filters
Robert Osazuwa Ness, Katie Matton, Hayden Helm +4
Large language models (LLM) have achieved impressive performance on medical question-answering benchmarks. However, high benchmark accuracy does not imply that the performance gene…