1 paper
Shunwen Bai, Ziping Ma, Chaoyang Zhang +4
The evaluation of LLM reasoning is moving from final-answer accuracy to process-level assessment, yet existing methods still fail to capture how models plan reasoning paths and all…