1 paper
Yanbiao Ma, Fei Luo, Linfeng Zhang +10
Reasoning in large language models is predominantly evaluated through labeled benchmarks, conflating task performance with the quality of internal inference. Here we study reasonin…