1 paper
Rezarta Islamaj, Robert Leaman, Joey Chan +13
Evaluating large language models (LLMs) in the biomedical domain requires benchmarks that can distinguish reasoning from pattern matching and remain discriminative as model capabil…