92 citations · 159 across the 12 of their papers we have counts for
1 paper · 1 filter
Shashwat Singh, Shauli Ravfogel, Jonathan Herzig +3
Language models often exhibit undesirable behavior, e.g., generating toxic or gender-biased text. In the case of neural language models, an encoding of the undesirable behavior is…