From the 1 of 5 linked papers with an AI index.
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
ReSyn: Autonomously Scaling Synthetic Environments for Reasoning Models
Andre He, Nathaniel Weir, Kaj Bostrom +4
Reinforcement learning with verifiable rewards (RLVR) has emerged as a promising approach for training reasoning language models (RLMs) by leveraging supervision from verifiers. Al…
cs.AI2025
VeriCoT: Neuro-symbolic Chain-of-Thought Validation via Logical Consistency Checks
Yu Feng, Nathaniel Weir, Kaj Bostrom +5
LLMs can perform multi-step reasoning through Chain-of-Thought (CoT), but they cannot reliably verify their own logic. Even when they reach correct answers, the underlying reasonin…