4 papers
Steering Beyond the Support: Adversarial Training on Unsupervised Jailbroken Activation Simulation
Luoyu Chen, Weiqi Wang, Zhiyi Tian +5
Jailbreak prompts can trigger harmful completions on aligned LLMs, In accordance, safety steering has been proposed: test-time activation interventions that steer jailbreak activat…
Ellipsoid Control: A White-list Jailbreak Defense via Benign Latent Modeling
Luoyu Chen, Weiqi Wang, Zhiyi Tian +3
Representation engineering (RepE) defenses have shown strong robustness against jailbreak attacks on large language models (LLMs). However, these methods fundamentally rely on blac…
Approximate Machine Unlearning through Manifold Representation Forgetting Guided by Self Mode Connectivity
Weiqi Wang, Zhiyi Tian, Chenhan Zhang +2
Machine unlearning is a fundamental mechanism that enforces the right to be forgotten. Existing unlearning studies that rely on label manipulation or task-gradient reversal often d…
EVE: Efficient Verification of Data Erasure through Customized Perturbation in Approximate Unlearning
Weiqi Wang, Zhiyi Tian, Chenhan Zhang +2
Verifying whether the machine unlearning process has been properly executed is critical but remains underexplored. Some existing approaches propose unlearning verification methods…