most citedTraining Convolutional Networks with Noisy Labels

514 citations · 514 across the 1 of their papers we have counts for

collaborators

10 papers

cs.CL20245 cited

Branch-Train-MiX: Mixing Expert LLMs into a Mixture-of-Experts LLM

Sainbayar Sukhbaatar, Olga Golovneva, Vasu Sharma +8

We investigate efficient methods for training Large Language Models (LLMs) to possess capabilities in multiple specialized domains, such as coding, math reasoning and world knowled…

cs.LG20243 cited

Teaching Large Language Models to Reason with Reinforcement Learning

Alex Havrilla, Yuqing Du, Sharath Chandra Raparthy +6

Reinforcement Learning from Human Feedback (\textbf{RLHF}) has emerged as a dominant approach for aligning LLM outputs with human preferences. Inspired by the success of RLHF, we s…

cs.CL20241 cited

Reverse Training to Nurse the Reversal Curse

Olga Golovneva, Zeyuan Allen-Zhu, Jason Weston +1

Large language models (LLMs) have a surprising failure: when trained on "A has a feature B", they do not generalize to "B is a feature of A", which is termed the Reversal Curse. Ev…

cs.AI20246 cited

Beyond A*: Better Planning with Transformers via Search Dynamics Bootstrapping

Lucas Lehnert, Sainbayar Sukhbaatar, DiJia Su +4

While Transformers have enabled tremendous progress in various application settings, such architectures still trail behind traditional symbolic planners for solving complex decisio…

cs.LG2023

A Data Source for Reasoning Embodied Agents

Jack Lanchantin, Sainbayar Sukhbaatar, Gabriel Synnaeve +3

Recent progress in using machine learning models for reasoning tasks has been driven by novel model architectures, large-scale pre-training protocols, and dedicated reasoning datas…

cs.CL20234 cited

Improving Open Language Models by Learning from Organic Interactions

Jing Xu, Da Ju, Joshua Lane +10

We present BlenderBot 3x, an update on the conversational model BlenderBot 3, which is now trained using organic conversation and feedback data from participating users of the syst…