241 citations · 1.2k across the 37 of their papers we have counts for
1 paper · 2 filters
Boxin Wang, Wei Ping, Chaowei Xiao +6
Pre-trained language models (LMs) are shown to easily generate toxic language. In this work, we systematically explore domain-adaptive training to reduce the toxicity of language m…