2 citations · 3 across the 8 of their papers we have counts for
1 paper · 1 filter
Ashutosh Baheti, Maarten Sap, Alan Ritter +1
Dialogue models trained on human conversations inadvertently learn to generate toxic responses. In addition to producing explicitly offensive utterances, these models can also impl…