9 citations · 9 across the 4 of their papers we have counts for
1 paper · 1 filter
Jiayu Liu, Zhenya Huang, Wei Dai +7
Although large language models (LLMs) show promise in solving complex mathematical tasks, existing evaluation paradigms rely solely on a coarse measure of overall answer accuracy,…