1 paper
Jianzhe Chai, Yu Zhe, Jun Sakuma
Benchmark-based evaluation is the de facto standard for comparing large language models (LLMs). However, its reliability is increasingly threatened by test set contamination, where…