1 paper · 1 filter
Zhiling Yan, Dingjie Song, Zhe Fang +4
The deployment of Large Language Models (LLMs) in high-stakes clinical settings demands rigorous and reliable evaluation. However, existing medical benchmarks remain static, suffer…