16 citations · 26 across the 8 of their papers we have counts for
1 paper · 1 filter
Xavier Suau, Pieter Delobelle, Katherine Metcalf +4
An important issue with Large Language Models (LLMs) is their undesired ability to generate toxic language. In this work, we show that the neurons responsible for toxicity can be d…