391 citations · 617 across the 4 of their papers we have counts for
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2022★ 391 cited
Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Yuntao Bai, Andy Jones, Kamal Ndousse +28
We apply preference modeling and reinforcement learning from human feedback (RLHF) to finetune language models to act as helpful and harmless assistants. We find this alignment tra…
cs.CL2021★ 130 cited
Understanding the Capabilities, Limitations, and Societal Impact of Large Language Models
Alex Tamkin, Miles Brundage, Jack Clark +1
On October 14th, 2020, researchers from OpenAI, the Stanford Institute for Human-Centered Artificial Intelligence, and other universities convened to discuss open research question…