11 papers
Soft Guidance Starts to Outperform CoT Prompting as LLMs Improve
Denys Pushkin, Albert Q. Jiang, Aryo Lotfi +3
Chain-of-Thought (CoT) prompting remains the standard baseline for evaluating models' reasoning abilities. Originally, this technique was introduced to elicit step-by-step reasonin…
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models
Matteo Santelmo, Xiuying Wei, Israa Fakih +5
Modern AI models achieve strong performance on many established benchmarks, yet they still fail on tasks that humans find almost trivial, such as manipulating a string or drawing a…
Goldilocks RL: Tuning Task Difficulty to Escape Sparse Rewards for Reasoning
Ilia Mahrooghi, Aryo Lotfi, Emmanuel Abbe
Reinforcement learning has emerged as a powerful paradigm for unlocking reasoning capabilities in language models. However, relying on sparse rewards makes this process highly samp…
LEAD: Breaking the No-Recovery Bottleneck in Long-Horizon Reasoning
Denys Pushkin, Emmanuel Abbe
Long-horizon execution in Large Language Models (LLMs) remains unstable even when high-level strategies are provided. Evaluating on controlled algorithmic puzzles, we demonstrate t…
RL for Reasoning by Adaptively Revealing Rationales
Mohammad Hossein Amani, Aryo Lotfi, Nicolas Mario Baldwin +4
Learning in the combinatorially large output space of sequence generation problems is challenging as providing expert demonstrations scales poorly with sequence length, and RL stru…
AbstRaL: Augmenting LLMs' Reasoning by Reinforcing Abstract Thinking
Silin Gao, Antoine Bosselut, Samy Bengio +1
Recent studies have shown that large language models (LLMs), especially smaller ones, often lack robustness in grade school math (GSM) reasoning. In particular, they tend to experi…