1 paper · 1 filter
Wentao Long, Yunfei Zhang, Chenyi Li +3
Formal theorem-proving benchmarks enable mechanically verifiable evaluation of mathematical reasoning in large language models. However, existing benchmarks mainly focus on Olympia…