1 paper · 1 filter
Zhizhang Fu, Yuancheng Gu, Chenkai Hu +2
Test-time scaling has enabled Large Language Models (LLMs) to tackle complex reasoning, yet the limitations of current Chain-of-Thought (CoT) evaluation obscures whether performanc…