1 citations · 1 across the 5 of their papers we have counts for
1 paper · 1 filter
Ricardo Dominguez-Olmedo, Florian E. Dorner, Moritz Hardt
We study a fundamental problem in the evaluation of large language models that we call training on the test task. Unlike wrongful practices like training on the test data, leakage,…