activity
20182022
most citedTraining a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

391 citations · 413 across the 2 of their papers we have counts for

collaborators

7 papers

cs.LG202222 cited

Scaling Laws and Interpretability of Learning from Repeated Data

Danny Hernandez, Tom Brown, Tom Conerly +15

Recent large language models have been trained on vast datasets, but also often on repeated data, either intentionally for the purpose of upweighting higher quality data, or uninte…

cs.CL2022391 cited

Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Yuntao Bai, Andy Jones, Kamal Ndousse +28

We apply preference modeling and reinforcement learning from human feedback (RLHF) to finetune language models to act as helpful and harmless assistants. We find this alignment tra…

cs.LG2019

Dota 2 with Large Scale Deep Reinforcement Learning

OpenAI, :, Christopher Berner +24

On April 13th, 2019, OpenAI Five became the first AI system to defeat the world champions at an esports game. The game of Dota 2 presents novel challenges for AI systems such as lo…

stat.ML2018

Discriminator Rejection Sampling

Samaneh Azadi, Catherine Olsson, Trevor Darrell +2

We propose a rejection sampling scheme using the discriminator of a GAN to approximately correct errors in the GAN generator distribution. We show that under quite strict assumptio…

stat.ML2018

Unrestricted Adversarial Examples

Tom B. Brown, Nicholas Carlini, Chiyuan Zhang +3

We introduce a two-player contest for evaluating the safety and robustness of machine learning systems, with a large prize pool. Unlike most prior work in ML robustness, which stud…

stat.ML2018

Skill Rating for Generative Models

Catherine Olsson, Surya Bhupatiraju, Tom Brown +2

We explore a new way to evaluate generative models using insights from evaluation of competitive games between human players. We show experimentally that tournaments between genera…