1 paper
Jixiang Hong, Quan Tu, Changyu Chen +3
Language models trained on large-scale corpus often generate content that is harmful, toxic, or contrary to human preferences, making their alignment with human values a critical c…