Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
FACT-E: Causality-Inspired Evaluation for Trustworthy Chain-of-Thought Reasoning
Yuxi Sun, Aoqi Zuo, Haotian Xie +3
Chain-of-Thought (CoT) prompting has improved LLM reasoning, but models often generate explanations that appear coherent while containing unfaithful intermediate steps. Existing se…
cs.AI2026
3D Instruction Ambiguity Detection
Jiayu Ding, Haoran Tang, Hongbo Jin +2
In safety-critical domains, linguistic ambiguity can have severe consequences; a vague command like "Pass me the vial" in a surgical setting could lead to catastrophic errors. Yet,…