1 paper
Narun Raman, Taylor Lundy, Thiago Amin +2
How should one judge whether a given large language model (LLM) can reliably perform economic reasoning? Most existing LLM benchmarks focus on specific applications and fail to pre…