2 papers
cs.CL2026
Controlled Reformulation Testing for Logical Consistency in Large Language Models
Alexander Gu, Alan Chen
Large language models (LLMs) frequently contradict themselves when the surface form of a logically equivalent question changes. We present a benchmark of 350 question families (1,7…
cs.AI2025
Solving Inequality Proofs with Large Language Models
Pan Lu, Jiayi Sheng, Luna Lyu +4
Inequality proving, crucial across diverse scientific and mathematical fields, tests advanced reasoning skills such as discovering tight bounds and strategic theorem application. T…