#reasoning evaluation

try —

4 papers match

cs.CL2026

Can LVLMs Uncover the Truth Behind Visual Illusions? An Analysis of Perceptual and Reasoning Capabilities

Liangjie Zhao, Jiaqing Lyu, Kexin Tang +5

The paper introduces IllusionReasoning, a benchmark that uses visual illusion images to jointly assess perception and reasoning abilities of large vision‑language models, revealing…

#visual illusions#vision-language models#reasoning evaluation#benchmark
cs.CR2026

A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models

Vivek Shukla, Varun Shukla, Atul +2

The paper proposes the Reasoning Answer Faithfulness Score (RAFS), a reference‑free metric that evaluates whether a large language model’s mathematical chain‑of‑thought trace is lo…

#large language models#chain of thought#reasoning evaluation#reference-free metrics
cs.CL2026

Controlled Reformulation Testing for Logical Consistency in Large Language Models

Alexander Gu, Alan Chen

The paper introduces CRTBench, a benchmark of 350 question families to test whether large language models give consistent answers across controlled reformulations such as contrapos…

#logical consistency#large language models#benchmark#reformulation testing
cs.CL2026

Filtered Reasoning Score: Evaluating Reasoning Quality on a Model's Most-Confident Traces

Manas Pathak, Xingyao Chen, Shuozhe Li +2

The paper introduces the Filtered Reasoning Score (FRS), a metric that evaluates the quality of reasoning traces from large language models by focusing on the most confident genera…

#large language models#reasoning evaluation#confidence filtering#benchmarking

One search, two signals: results blend meaning (embedding similarity, so papers that never use your words still surface) with keyword matches on titles, abstracts and summaries. Free, no sign-in needed.