1 paper · 1 filter
Kun Wang, Reinhard Heckel
Direct evaluation of LLMs on benchmarks can be misleading because comparatively strong performance may reflect task familiarity rather than capability. The train-before-test approa…