1 paper · 1 filter
Xia Yang, Xuanyi Zhang, Hao Hu +1
Large language models now achieve high final-answer accuracy on mathematical reasoning benchmarks, but accuracy alone does not capture reasoning flexibility. We introduce a strateg…