1 paper · 1 filter
Alexandra Abbas, Nora Petrova, Helios Ael Lyons +1
Recent work has shown that language models' refusal behavior is primarily encoded in a single direction in their latent space, making it vulnerable to targeted attacks. Although La…