391 citations · 413 across the 2 of their papers we have counts for
7 papers
Scaling Laws and Interpretability of Learning from Repeated Data
Danny Hernandez, Tom Brown, Tom Conerly +15
Recent large language models have been trained on vast datasets, but also often on repeated data, either intentionally for the purpose of upweighting higher quality data, or uninte…
Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Yuntao Bai, Andy Jones, Kamal Ndousse +28
We apply preference modeling and reinforcement learning from human feedback (RLHF) to finetune language models to act as helpful and harmless assistants. We find this alignment tra…
Dota 2 with Large Scale Deep Reinforcement Learning
OpenAI, :, Christopher Berner +24
On April 13th, 2019, OpenAI Five became the first AI system to defeat the world champions at an esports game. The game of Dota 2 presents novel challenges for AI systems such as lo…
Discriminator Rejection Sampling
Samaneh Azadi, Catherine Olsson, Trevor Darrell +2
We propose a rejection sampling scheme using the discriminator of a GAN to approximately correct errors in the GAN generator distribution. We show that under quite strict assumptio…
Unrestricted Adversarial Examples
Tom B. Brown, Nicholas Carlini, Chiyuan Zhang +3
We introduce a two-player contest for evaluating the safety and robustness of machine learning systems, with a large prize pool. Unlike most prior work in ML robustness, which stud…
Skill Rating for Generative Models
Catherine Olsson, Surya Bhupatiraju, Tom Brown +2
We explore a new way to evaluate generative models using insights from evaluation of competitive games between human players. We show experimentally that tournaments between genera…