20 citations · 53 across the 17 of their papers we have counts for
1 paper · 2 filters
Stephen Miner, Yoshiki Takashima, Simeng Han +4
Benchmarks are critical for measuring Large Language Model (LLM) reasoning capabilities. Some benchmarks have even become the de facto indicator of such capabilities. However, as L…