1 paper
Batu Guan, Xiao Wu, Yuanyuan Yuan +1
In this paper, we tackle a critical challenge in model evaluation: how to keep code benchmarks useful when models might have already seen them during training. We introduce a novel…