48 citations · 249 across the 35 of their papers we have counts for
16 papers · 1 filter
SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks
Yifei Zhou, Song Jiang, Yuandong Tian +4
Large language model (LLM) agents need to perform multi-turn interactions in real-world tasks. However, existing multi-turn RL algorithms for optimizing LLM agents fail to perform…
Teaching Large Language Models to Reason with Reinforcement Learning
Alex Havrilla, Yuqing Du, Sharath Chandra Raparthy +6
Reinforcement Learning from Human Feedback (\textbf{RLHF}) has emerged as a dominant approach for aligning LLM outputs with human preferences. Inspired by the success of RLHF, we s…
A Data Source for Reasoning Embodied Agents
Jack Lanchantin, Sainbayar Sukhbaatar, Gabriel Synnaeve +3
Recent progress in using machine learning models for reasoning tasks has been driven by novel model architectures, large-scale pre-training protocols, and dedicated reasoning datas…
Large Language Model Programs
Imanol Schlag, Sainbayar Sukhbaatar, Asli Celikyilmaz +4
In recent years, large pre-trained language models (LLMs) have demonstrated the ability to follow instructions and perform novel tasks from a few examples. The possibility to param…
Learning to Reason and Memorize with Self-Notes
Jack Lanchantin, Shubham Toshniwal, Jason Weston +2
Large language models have been shown to struggle with multi-step reasoning, and do not retain previous reasoning steps for future use. We propose a simple method for solving both…
Walk the Random Walk: Learning to Discover and Reach Goals Without Supervision
Lina Mezghani, Sainbayar Sukhbaatar, Piotr Bojanowski +1
Learning a diverse set of skills by interacting with an environment without any external supervision is an important challenge. In particular, obtaining a goal-conditioned agent th…