1 paper · 1 filter
Dongxu Zhang, Yiding Sun, Zihao Guo +5
Reasoning failures in large language models (LLMs) are usually evaluated from final answers, but a wrong answer does not reveal why the model failed. The same incorrect output may…