14 citations · 14 across the 2 of their papers we have counts for
1 paper · 1 filter
Tomohiro Sawada, Daniel Paleka, Alexander Havrilla +6
Large Language Models (LLMs) have demonstrated remarkable performance on various quantitative reasoning and knowledge benchmarks. However, many of these benchmarks are losing utili…