7 papers
Soft Guidance Starts to Outperform CoT Prompting as LLMs Improve
Denys Pushkin, Albert Q. Jiang, Aryo Lotfi +3
Chain-of-Thought (CoT) prompting remains the standard baseline for evaluating models' reasoning abilities. Originally, this technique was introduced to elicit step-by-step reasonin…
Goldilocks RL: Tuning Task Difficulty to Escape Sparse Rewards for Reasoning
Ilia Mahrooghi, Aryo Lotfi, Emmanuel Abbe
Reinforcement learning has emerged as a powerful paradigm for unlocking reasoning capabilities in language models. However, relying on sparse rewards makes this process highly samp…
RL for Reasoning by Adaptively Revealing Rationales
Mohammad Hossein Amani, Aryo Lotfi, Nicolas Mario Baldwin +4
Learning in the combinatorially large output space of sequence generation problems is challenging as providing expert demonstrations scales poorly with sequence length, and RL stru…
To Infinity and Beyond: Tool-Use Unlocks Length Generalization in State Space Models
Eran Malach, Omid Saremi, Sinead Williamson +5
State Space Models (SSMs) have become the leading alternative to Transformers for sequence modeling. Their primary advantage is efficiency in long-context and long-form generation,…
Chain-of-Sketch: Enabling Global Visual Reasoning
Aryo Lotfi, Enrico Fini, Samy Bengio +2
Modern vision models have achieved remarkable success in benchmarks where local features provide critical information about the target. There is now a growing interest in tackling…
Generalization on the Unseen, Logic Reasoning and Degree Curriculum
Emmanuel Abbe, Samy Bengio, Aryo Lotfi +1
This paper considers the learning of logical (Boolean) functions with a focus on the generalization on the unseen (GOTU) setting, a strong case of out-of-distribution generalizatio…