Showing cs.CLShow all
3 papers · 1 filter
cs.CL2025
Semi-structured LLM Reasoners Can Be Rigorously Audited
Jixuan Leng, Cassandra A. Cohen, Zhixian Zhang +2
Although Large Language Models (LLMs) have become capable reasoners, the problem of faithfulness persists: their reasoning can contain errors and omissions that are difficult to de…
cs.CL2025
CrossWordBench: Evaluating the Reasoning Capabilities of LLMs and LVLMs with Controllable Puzzle Generation
Jixuan Leng, Chengsong Huang, Langlin Huang +4
Existing reasoning evaluation frameworks for Large Language Models (LLMs) and Large Vision-Language Models (LVLMs) predominantly assess either text-based reasoning or vision-langua…
cs.CL2024
Watch Your Steps: Observable and Modular Chains of Thought
Cassandra A. Cohen, William W. Cohen
We propose a variant of chain of thought (CoT) prompting called Program Trace Prompting that makes explanations more observable while preserving the power, generality and flexibili…