3 papers
cs.AI2026
Localizing and Correcting Errors for LLM-based Planners
Aditya Kumar, William W. Cohen
Large language models (LLMs) have demonstrated strong reasoning capabilities on math and coding, but frequently fail on symbolic classical planning tasks. Our studies, as well as p…
cs.CL2025
Semi-structured LLM Reasoners Can Be Rigorously Audited
Jixuan Leng, Cassandra A. Cohen, Zhixian Zhang +2
Although Large Language Models (LLMs) have become capable reasoners, the problem of faithfulness persists: their reasoning can contain errors and omissions that are difficult to de…
cs.CL2025
CrossWordBench: Evaluating the Reasoning Capabilities of LLMs and LVLMs with Controllable Puzzle Generation
Jixuan Leng, Chengsong Huang, Langlin Huang +4
Existing reasoning evaluation frameworks for Large Language Models (LLMs) and Large Vision-Language Models (LVLMs) predominantly assess either text-based reasoning or vision-langua…