8 citations · 8 across the 1 of their papers we have counts for
1 paper
Sandipan Kundu, Yuntao Bai, Saurav Kadavath +33
Human feedback can prevent overtly harmful utterances in conversational models, but may not automatically mitigate subtle problematic behaviors such as a stated desire for self-pre…