Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Soft Guidance Starts to Outperform CoT Prompting as LLMs Improve
Denys Pushkin, Albert Q. Jiang, Aryo Lotfi +3
Chain-of-Thought (CoT) prompting remains the standard baseline for evaluating models' reasoning abilities. Originally, this technique was introduced to elicit step-by-step reasonin…
cs.AI2026
LEAD: Breaking the No-Recovery Bottleneck in Long-Horizon Reasoning
Denys Pushkin, Emmanuel Abbe
Long-horizon execution in Large Language Models (LLMs) remains unstable even when high-level strategies are provided. Evaluating on controlled algorithmic puzzles, we demonstrate t…