391 citations · 530 across the 3 of their papers we have counts for
3 papers
Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Yuntao Bai, Andy Jones, Kamal Ndousse +28
We apply preference modeling and reinforcement learning from human feedback (RLHF) to finetune language models to act as helpful and harmless assistants. We find this alignment tra…
The AI Index 2021 Annual Report
Daniel Zhang, Saurabh Mishra, Erik Brynjolfsson +10
Welcome to the fourth edition of the AI Index Report. This year we significantly expanded the amount of data available in the report, worked with a broader set of external organiza…
Understanding the Capabilities, Limitations, and Societal Impact of Large Language Models
Alex Tamkin, Miles Brundage, Jack Clark +1
On October 14th, 2020, researchers from OpenAI, the Stanford Institute for Human-Centered Artificial Intelligence, and other universities convened to discuss open research question…