Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Latent-space Attacks for Refusal Evasion in Language Models
Giorgio Piras, Raffaele Mura, Fabio Brau +4
Safety-aligned language models are trained to refuse harmful requests, yet refusal behavior can be suppressed by steering their internal representations. Existing methods do so by…
cs.AI2025
SOM Directions are Better than One: Multi-Directional Refusal Suppression in Language Models
Giorgio Piras, Raffaele Mura, Fabio Brau +3
Refusal refers to the functional behavior enabling safety-aligned language models to reject harmful or unethical prompts. Following the growing scientific interest in mechanistic i…