1 paper · 1 filter
Natasha Jaques, Judy Hanwen Shen, Asma Ghandeharioun +5
How can we train a dialog model to produce better conversations by learning from human feedback, without the risk of humans teaching it harmful chat behaviors? We start by hosting…