1 paper · 1 filter
Yebo Peng, Yaoming Li, Zixiang Liu +6
Evaluating the mathematical capability of Large Language Models (LLMs) is a critical yet challenging frontier. Existing benchmarks fall short, particularly for proof-centric proble…