1 paper · 1 filter
Shi-Yu Tian, Zhi Zhou, Wei Dong +5
Mathematical reasoning has long been a key benchmark for evaluating large language models. Although substantial progress has been made on math word problems, the need for reasoning…