4 papers
ReEfBench: Quantifying the Reasoning Efficiency of LLMs
Zhizhang Fu, Yuancheng Gu, Chenkai Hu +2
Test-time scaling has enabled Large Language Models (LLMs) to tackle complex reasoning, yet the limitations of current Chain-of-Thought (CoT) evaluation obscures whether performanc…
Correlation or Causation: Analyzing the Causal Structures of LLM and LRM Reasoning Process
Zhizhang FU, Guangsheng Bao, Hongbo Zhang +2
LLMs suffer from critical reasoning issues such as unfaithfulness, bias, and inconsistency, since they lack robust causal underpinnings and may rely on superficial correlations rat…
Evaluating the Logical Reasoning Abilities of Large Reasoning Models
Hanmeng Liu, Yiran Ding, Zhizhang Fu +3
Large reasoning models, often post-trained on long chain-of-thought (long CoT) data with reinforcement learning, achieve state-of-the-art performance on mathematical, coding, and d…
Logical Reasoning in Large Language Models: A Survey
Hanmeng Liu, Zhizhang Fu, Mengru Ding +4
With the emergence of advanced reasoning models like OpenAI o3 and DeepSeek-R1, large language models (LLMs) have demonstrated remarkable reasoning capabilities. However, their abi…