4 papers · 1 filter
QuickScope: Certifying Hard Questions in Dynamic LLM Benchmarks
Taylor Lundy, Narun K. Raman, Kevin Leyton-Brown
LLM benchmarks are increasingly dynamic: instead of containing a fixed set of questions, they define templates and parameters that can generate an effectively unlimited number of q…
Reasoning Models are Test Exploiters: Rethinking Multiple-Choice
Narun Raman, Taylor Lundy, Kevin Leyton-Brown
When evaluating Large Language Models (LLMs) in question answering domains, it is common to ask the model to choose among a fixed set of choices (so-called multiple-choice question…
STEER-ME: Assessing the Microeconomic Reasoning of Large Language Models
Narun Raman, Taylor Lundy, Thiago Amin +2
How should one judge whether a given large language model (LLM) can reliably perform economic reasoning? Most existing LLM benchmarks focus on specific applications and fail to pre…
STEER: Assessing the Economic Rationality of Large Language Models
Narun Raman, Taylor Lundy, Samuel Amouyal +3
There is increasing interest in using LLMs as decision-making "agents." Doing so includes many degrees of freedom: which model should be used; how should it be prompted; should it…