1 paper
Xia Yang, Xuanyi Zhang, Hao Hu +1
Large language models now achieve high final-answer accuracy on mathematical reasoning benchmarks, but accuracy alone does not capture reasoning flexibility. We introduce a strateg…