170 citations · 342 across the 3 of their papers we have counts for
3 papers · 1 filter
The Capacity for Moral Self-Correction in Large Language Models
Deep Ganguli, Amanda Askell, Nicholas Schiefer +46
We test the hypothesis that language models trained with reinforcement learning from human feedback (RLHF) have the capability to "morally self-correct" -- to avoid producing harmf…
Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
Deep Ganguli, Liane Lovitt, Jackson Kernion +33
We describe our early efforts to red team language models in order to simultaneously discover, measure, and attempt to reduce their potentially harmful outputs. We make three main…
Language Models (Mostly) Know What They Know
Saurav Kadavath, Tom Conerly, Amanda Askell +33
We study whether language models can evaluate the validity of their own claims and predict which questions they will be able to answer correctly. We first show that larger models a…