1 paper · 1 filter
Thanni Adewuyi, Anuoluwa Sotome, Samuel Okoko +6
Large language models achieve strong scores on medical benchmarks, yet these benchmarks evaluate each question in isolation, providing no measure of whether a system can distinguis…