1 paper
Taylor Lundy, Narun K. Raman, Kevin Leyton-Brown
LLM benchmarks are increasingly dynamic: instead of containing a fixed set of questions, they define templates and parameters that can generate an effectively unlimited number of q…