activity
20192022
most citedTraining a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

391 citations · 391 across the 3 of their papers we have counts for

collaborators

6 papers

cs.LG2022

DeepChrome 2.0: Investigating and Improving Architectures, Visualizations, & Experiments

Saurav Kadavath, Samuel Paradis, Jacob Yeung

Histone modifications play a critical role in gene regulation. Consequently, predicting gene expression from histone modification signals is a highly motivated problem in epigeneti…

cs.CL2022391 cited

Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Yuntao Bai, Andy Jones, Kamal Ndousse +28

We apply preference modeling and reinforcement learning from human feedback (RLHF) to finetune language models to act as helpful and harmless assistants. We find this alignment tra…

cs.LG2021

Pretraining & Reinforcement Learning: Sharpening the Axe Before Cutting the Tree

Saurav Kadavath, Samuel Paradis, Brian Yao

Pretraining is a common technique in deep learning for increasing performance and reducing training time, with promising experimental results in deep reinforcement learning (RL). H…

cs.SE2021

Measuring Coding Challenge Competence With APPS

Dan Hendrycks, Steven Basart, Saurav Kadavath +8

While programming is one of the most broadly applicable skills in modern society, modern machine learning models still cannot code solutions to basic problems. Despite its importan…

cs.LG2021

Measuring Mathematical Problem Solving With the MATH Dataset

Dan Hendrycks, Collin Burns, Saurav Kadavath +5

Many intellectual endeavors require mathematical problem solving, but this skill remains beyond the capabilities of computers. To measure this ability in machine learning models, w…

cs.LG2019

Using Self-Supervised Learning Can Improve Model Robustness and Uncertainty

Dan Hendrycks, Mantas Mazeika, Saurav Kadavath +1

Self-supervision provides effective representations for downstream tasks without requiring labels. However, existing approaches lag behind fully supervised training and are often n…