3 papers
cs.CL2025
RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation
Xinnuo Xu, Rachel Lawrence, Kshitij Dubey +7
Recent Large Language Models (LLMs) have reported high accuracy on reasoning benchmarks. However, it is still unclear whether the observed results arise from true reasoning or from…
cs.CL2025
Compositional Causal Reasoning Evaluation in Language Models
Jacqueline R. M. A. Maasch, Alihan Hüyük, Xinnuo Xu +2
Causal reasoning and compositional reasoning are two core aspirations in AI. Measuring the extent of these behaviors requires principled evaluation methods. We explore a unified pe…
cs.CL2025
Reasoning Elicitation in Language Models via Counterfactual Feedback
Alihan Hüyük, Xinnuo Xu, Jacqueline Maasch +2
Despite the increasing effectiveness of language models, their reasoning capabilities remain underdeveloped. In particular, causal reasoning through counterfactual question answeri…