1 paper
Dominic Okonkwo, Magnus Hodgson, Temitope I. David +1
Frontier language models are increasingly evaluated on biomedical benchmarks, but two problems undermine most published evaluations: legacy benchmarks are near-saturated, and open-…