1 paper · 1 filter
Tianxi Gao, Yufan Cai, Yusi Yuan +1
Large language models (LLMs) achieve promising performance, yet their ability to reason remains poorly understood. Existing evaluations largely emphasize task-level accuracy, often…