1 paper
Xavier Suau, Pieter Delobelle, Katherine Metcalf +4
An important issue with Large Language Models (LLMs) is their undesired ability to generate toxic language. In this work, we show that the neurons responsible for toxicity can be d…