1 citations · 1 across the 8 of their papers we have counts for
1 paper · 1 filter
Shu-Xun Yang, Cunxiang Wang, Yidong Wang +3
Evaluating mathematical capabilities is critical for assessing the overall performance of large language models (LLMs). However, existing evaluation methods often focus solely on f…