1 paper
Shi-Yu Tian, Zhi Zhou, Wei Dong +5
Mathematical reasoning has long been a key benchmark for evaluating large language models. Although substantial progress has been made on math word problems, the need for reasoning…