2 papers
cs.CR2026
Steering Beyond the Support: Adversarial Training on Unsupervised Jailbroken Activation Simulation
Luoyu Chen, Weiqi Wang, Zhiyi Tian +5
Jailbreak prompts can trigger harmful completions on aligned LLMs, In accordance, safety steering has been proposed: test-time activation interventions that steer jailbreak activat…
cs.CR2026
Ellipsoid Control: A White-list Jailbreak Defense via Benign Latent Modeling
Luoyu Chen, Weiqi Wang, Zhiyi Tian +3
Representation engineering (RepE) defenses have shown strong robustness against jailbreak attacks on large language models (LLMs). However, these methods fundamentally rely on blac…