activity
20232026
most citedChain-of-Symbol Prompting Elicits Planning in Large Langauge Models

7 citations · 14 across the 12 of their papers we have counts for

collaborators
Showing cs.AIShow all

6 papers · 1 filter

cs.AI2026

ReEfBench: Quantifying the Reasoning Efficiency of LLMs

Zhizhang Fu, Yuancheng Gu, Chenkai Hu +2

Test-time scaling has enabled Large Language Models (LLMs) to tackle complex reasoning, yet the limitations of current Chain-of-Thought (CoT) evaluation obscures whether performanc…

cs.AI2025

LOGicalThought: Logic-Based Ontological Grounding of LLMs for High-Assurance Reasoning

Navapat Nananukul, Yue Zhang, Ryan Lee +5

High-assurance reasoning, particularly in critical domains such as law and medicine, requires conclusions that are accurate, verifiable, and explicitly grounded in evidence. This r…

cs.AI2025

Correlation or Causation: Analyzing the Causal Structures of LLM and LRM Reasoning Process

Zhizhang FU, Guangsheng Bao, Hongbo Zhang +2

LLMs suffer from critical reasoning issues such as unfaithfulness, bias, and inconsistency, since they lack robust causal underpinnings and may rely on superficial correlations rat…

cs.AI2025

GAUSS: Benchmarking Structured Mathematical Skills for Large Language Models

Yue Zhang, Jiaxin Zhang, Qiuyu Ren +5

We introduce \textbf{GAUSS} (\textbf{G}eneral \textbf{A}ssessment of \textbf{U}nderlying \textbf{S}tructured \textbf{S}kills in Mathematics), a benchmark that evaluates LLMs' mathe…

cs.AI2025

Evaluating the Logical Reasoning Abilities of Large Reasoning Models

Hanmeng Liu, Yiran Ding, Zhizhang Fu +3

Large reasoning models, often post-trained on long chain-of-thought (long CoT) data with reinforcement learning, achieve state-of-the-art performance on mathematical, coding, and d…

cs.AI2025★ 5 cited

Logical Reasoning in Large Language Models: A Survey

Hanmeng Liu, Zhizhang Fu, Mengru Ding +4

With the emergence of advanced reasoning models like OpenAI o3 and DeepSeek-R1, large language models (LLMs) have demonstrated remarkable reasoning capabilities. However, their abi…