1 paper
Wei Shao, Yihang Wang, Gaoyu Zhu +4
Existing detoxification methods for large language models mainly focus on post-training stage or inference time, while few tackle the source of toxicity, namely, the dataset itself…