benchmark 1large language models 1logical consistency 1reasoning evaluation 1reformulation testing 1
From the 1 of 3 linked papers with an AI index.
3 papers
cs.CL2026
Controlled Reformulation Testing for Logical Consistency in Large Language Models
Alexander Gu, Alan Chen
The paper introduces CRTBench, a benchmark of 350 question families to test whether large language models give consistent answers across controlled reformulations such as contrapos…
cs.AI2026
Investigating Advanced Reasoning of Large Language Models via Black-Box Environment Interaction
Congchi Yin, Tianyi Wu, Yankai Shu +5
Existing tasks fall short in evaluating reasoning ability of Large Language Models (LLMs) in an interactive, unknown environment. This deficiency leads to the isolated assessment o…
cs.AI2025
Solving Inequality Proofs with Large Language Models
Pan Lu, Jiayi Sheng, Luna Lyu +4
Inequality proving, crucial across diverse scientific and mathematical fields, tests advanced reasoning skills such as discovering tight bounds and strategic theorem application. T…