1 paper · 1 filter
Xidong Wang, Shuqi Guo, Yue Shen +6
The reliability of medical LLM evaluation is critically undermined by data contamination and knowledge obsolescence, leading to inflated scores on static benchmarks. To address the…