1 citations · 1 across the 3 of their papers we have counts for
1 paper · 1 filter
Stephen Miner, Yoshiki Takashima, Simeng Han +4
Benchmarks are critical for measuring Large Language Model (LLM) reasoning capabilities. Some benchmarks have even become the de facto indicator of such capabilities. However, as L…