Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
How Robust Are Automated Fact-Checking Systems? A Cross-Benchmark Evaluation
Aida Usmanova, Zangir Iklassov, Markus Leippold +1
Automated fact-checking (AFC) systems retrieve evidence and predict claim veracity, yet evaluations omit simple baselines, systems are developed for a single benchmark and cannot b…
cs.AI2026
SymStep: Symbolic Step Verification for Logical Reasoning
Aida Usmanova, Rui Gao, Dilshod Azizov +2
Chain-of-thought (CoT) prompting can fail severely on constraint-dense logical reasoning tasks, where unverified errors accumulate silently across steps. We introduce SymStep: an L…