1 paper · 1 filter
Zenghao Duan, Zhiyi Yin, Zhichao Shi +6
This paper investigates the underlying mechanisms of toxicity generation in Large Language Models (LLMs) and proposes an effective detoxification approach. Prior work typically con…