1 paper · 1 filter
Zijian Chen, Wenjun Zhang, Guangtao Zhai
The potential data contamination issue in contemporary large language models (LLMs) benchmarks presents a fundamental challenge to establishing trustworthy evaluation frameworks. M…