5 papers
QuickScope: Certifying Hard Questions in Dynamic LLM Benchmarks
Taylor Lundy, Narun K. Raman, Kevin Leyton-Brown
LLM benchmarks are increasingly dynamic: instead of containing a fixed set of questions, they define templates and parameters that can generate an effectively unlimited number of q…
Reasoning Models are Test Exploiters: Rethinking Multiple-Choice
Narun Raman, Taylor Lundy, Kevin Leyton-Brown
When evaluating Large Language Models (LLMs) in question answering domains, it is common to ask the model to choose among a fixed set of choices (so-called multiple-choice question…
NFTs as a Data-Rich Test Bed: Conspicuous Consumption and its Determinants
Taylor Lundy, Narun Raman, Scott Duke Kominers +1
Conspicuous consumption occurs when a consumer derives value from a good based on its social meaning as a signal of wealth, taste, and/or community affiliation. Common conspicuous…
STEER-ME: Assessing the Microeconomic Reasoning of Large Language Models
Narun Raman, Taylor Lundy, Thiago Amin +2
How should one judge whether a given large language model (LLM) can reliably perform economic reasoning? Most existing LLM benchmarks focus on specific applications and fail to pre…
Multidimensional Bayesian Utility Maximization: Tight Approximations to Welfare
Kira Goldner, Taylor Lundy
We initiate the study of multidimensional Bayesian utility maximization, focusing on the unit-demand setting where values are i.i.d. across both items and buyers. The seminal resul…