4 papers · 1 filter
Soft Guidance Starts to Outperform CoT Prompting as LLMs Improve
Denys Pushkin, Albert Q. Jiang, Aryo Lotfi +3
Chain-of-Thought (CoT) prompting remains the standard baseline for evaluating models' reasoning abilities. Originally, this technique was introduced to elicit step-by-step reasonin…
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models
Matteo Santelmo, Xiuying Wei, Israa Fakih +5
Modern AI models achieve strong performance on many established benchmarks, yet they still fail on tasks that humans find almost trivial, such as manipulating a string or drawing a…
LEAD: Breaking the No-Recovery Bottleneck in Long-Horizon Reasoning
Denys Pushkin, Emmanuel Abbe
Long-horizon execution in Large Language Models (LLMs) remains unstable even when high-level strategies are provided. Evaluating on controlled algorithmic puzzles, we demonstrate t…
Algorithm Discovery With LLMs: Evolutionary Search Meets Reinforcement Learning
Anja Surina, Amin Mansouri, Lars Quaedvlieg +4
Discovering efficient algorithms for solving complex problems has been an outstanding challenge in mathematics and computer science, requiring substantial human expertise over the…