254 citations · 600 across the 38 of their papers we have counts for
1 paper · 2 filters
Nazanin Mohammadi Sepahvand, Eleni Triantafillou, Hugo Larochelle +3
Large language models (LLMs) trained on webscale data can produce toxic outputs, raising concerns for safe deployment. Prior defenses, based on applications of DPO, NPO, and simila…