Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
LinEAS: End-to-end Learning of Activation Steering with a Distributional Loss
Pau Rodriguez, Michal Klein, Eleonora Gualdoni +5
The growing use of generative models in daily life calls for efficient mechanisms to control their generation, to e.g., produce safe content or provide users with tools to explore…
cs.CL2024
Whispering Experts: Neural Interventions for Toxicity Mitigation in Language Models
Xavier Suau, Pieter Delobelle, Katherine Metcalf +4
An important issue with Large Language Models (LLMs) is their undesired ability to generate toxic language. In this work, we show that the neurons responsible for toxicity can be d…