1 paper
Ari Spiesberger, Juan J. Vazquez, Nicky Pochinkov +4
If LLM training data is polluted with benchmark test data, then benchmark performance gives biased estimates of out-of-distribution (OOD) generalization. Typical decontamination fi…