30 citations · 32 across the 13 of their papers we have counts for
1 paper · 1 filter
Stephen Miner, Yoshiki Takashima, Simeng Han +4
Benchmarks are critical for measuring Large Language Model (LLM) reasoning capabilities. Some benchmarks have even become the de facto indicator of such capabilities. However, as L…