1 paper
Dongxu Zhang, Yiding Sun, Zihao Guo +5
Reasoning failures in large language models (LLMs) are usually evaluated from final answers, but a wrong answer does not reveal why the model failed. The same incorrect output may…