400 citations · 553 across the 5 of their papers we have counts for
5 papers
Self-Taught Evaluators
Tianlu Wang, Ilia Kulikov, Olga Golovneva +7
Model-based evaluation is at the heart of successful model development -- as a reward model for training, and as a replacement for human evaluation. To train such evaluators, the s…
Teaching Large Language Models to Reason with Reinforcement Learning
Alex Havrilla, Yuqing Du, Sharath Chandra Raparthy +6
Reinforcement Learning from Human Feedback (\textbf{RLHF}) has emerged as a dominant approach for aligning LLM outputs with human preferences. Inspired by the success of RLHF, we s…
Shepherd: A Critic for Language Model Generation
Tianlu Wang, Ping Yu, Xiaoqing Ellen Tan +7
As large language models improve, there is increasing interest in techniques that leverage these models' capabilities to refine their own outputs. In this work, we introduce Shephe…
Augmented Language Models: a Survey
Grégoire Mialon, Roberto Dessì, Maria Lomeli +10
This survey reviews works in which language models (LMs) are augmented with reasoning skills and the ability to use tools. The former is defined as decomposing a potentially comple…
Toolformer: Language Models Can Teach Themselves to Use Tools
Timo Schick, Jane Dwivedi-Yu, Roberto Dessì +5
Language models (LMs) exhibit remarkable abilities to solve new tasks from just a few examples or textual instructions, especially at scale. They also, paradoxically, struggle with…