1 paper
Giorgio Piras, Raffaele Mura, Fabio Brau +4
Safety-aligned language models are trained to refuse harmful requests, yet refusal behavior can be suppressed by steering their internal representations. Existing methods do so by…