1 paper
Xiaodong Yu, Ben Zhou, Hao Cheng +1
Existing math datasets evaluate the reasoning abilities of large language models (LLMs) by either using the final answer or the intermediate reasoning steps derived from static exa…