1 paper
Bao Nguyen, Binh Nguyen, Duy Nguyen +1
Language models, while capable of generating remarkably coherent and seemingly accurate text, can occasionally produce undesirable content, including harmful or toxic outputs. In t…