53 citations · 92 across the 2 of their papers we have counts for
1 paper · 1 filter
Deep Ganguli, Amanda Askell, Nicholas Schiefer +46
We test the hypothesis that language models trained with reinforcement learning from human feedback (RLHF) have the capability to "morally self-correct" -- to avoid producing harmf…