activity
20132025
most citedHash Layers For Large Sparse Models

48 citations · 249 across the 35 of their papers we have counts for

collaborators
Showing cs.LGShow all

16 papers · 1 filter

cs.LG2025

SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Yifei Zhou, Song Jiang, Yuandong Tian +4

Large language model (LLM) agents need to perform multi-turn interactions in real-world tasks. However, existing multi-turn RL algorithms for optimizing LLM agents fail to perform…

cs.LG2024★ 3 cited

Teaching Large Language Models to Reason with Reinforcement Learning

Alex Havrilla, Yuqing Du, Sharath Chandra Raparthy +6

Reinforcement Learning from Human Feedback (\textbf{RLHF}) has emerged as a dominant approach for aligning LLM outputs with human preferences. Inspired by the success of RLHF, we s…

cs.LG2023

A Data Source for Reasoning Embodied Agents

Jack Lanchantin, Sainbayar Sukhbaatar, Gabriel Synnaeve +3

Recent progress in using machine learning models for reasoning tasks has been driven by novel model architectures, large-scale pre-training protocols, and dedicated reasoning datas…

cs.LG2023★ 5 cited

Large Language Model Programs

Imanol Schlag, Sainbayar Sukhbaatar, Asli Celikyilmaz +4

In recent years, large pre-trained language models (LLMs) have demonstrated the ability to follow instructions and perform novel tasks from a few examples. The possibility to param…

cs.LG2023★ 3 cited

Learning to Reason and Memorize with Self-Notes

Jack Lanchantin, Shubham Toshniwal, Jason Weston +2

Large language models have been shown to struggle with multi-step reasoning, and do not retain previous reasoning steps for future use. We propose a simple method for solving both…

cs.LG2022

Walk the Random Walk: Learning to Discover and Reach Goals Without Supervision

Lina Mezghani, Sainbayar Sukhbaatar, Piotr Bojanowski +1

Learning a diverse set of skills by interacting with an environment without any external supervision is an important challenge. In particular, obtaining a goal-conditioned agent th…