3 citations · 6 across the 3 of their papers we have counts for
1 paper · 1 filter
Zhaowei Zhu, Jialu Wang, Hao Cheng +1
Language models have shown promise in various tasks but can be affected by undesired data during training, fine-tuning, or alignment. For example, if some unsafe conversations are…