1 paper
Xingyao Xiao, Yihong Cheng
Benchmark contamination, the leakage of test items into training data, is widely described as a threat to the reliability of large language model (LLM) leaderboards. We argue that…