2 papers
cs.SE2026
Neither Layer Alone: Epistemic Integrity Requires Hierarchical Joint Design for Long-Running AI Agents
Zhihong Shen
Long-running AI agents fail not only when inference fails or tools are underspecified, but when independently evolving model and harness layers change the semantics of belief, capa…
cs.AI2026
Reasoning Fails Where Step Flow Breaks
Xiaoyu Xu, Yulan Pan, Xiaosong Yuan +4
Large reasoning models (LRMs) that generate long chains of thought now perform well on multi-step math, science, and coding tasks. However, their behavior is still unstable and har…