1 paper · 1 filter
Taylor Lundy, Narun K. Raman, Kevin Leyton-Brown
LLM benchmarks are increasingly dynamic: instead of containing a fixed set of questions, they define templates and parameters that can generate an effectively unlimited number of q…