10 papers
Steering Beyond the Support: Adversarial Training on Unsupervised Jailbroken Activation Simulation
Luoyu Chen, Weiqi Wang, Zhiyi Tian +5
Jailbreak prompts can trigger harmful completions on aligned LLMs, In accordance, safety steering has been proposed: test-time activation interventions that steer jailbreak activat…
Ellipsoid Control: A White-list Jailbreak Defense via Benign Latent Modeling
Luoyu Chen, Weiqi Wang, Zhiyi Tian +3
Representation engineering (RepE) defenses have shown strong robustness against jailbreak attacks on large language models (LLMs). However, these methods fundamentally rely on blac…
Approximate Machine Unlearning through Manifold Representation Forgetting Guided by Self Mode Connectivity
Weiqi Wang, Zhiyi Tian, Chenhan Zhang +2
Machine unlearning is a fundamental mechanism that enforces the right to be forgotten. Existing unlearning studies that rely on label manipulation or task-gradient reversal often d…
GeoIB: Geometry-Aware Information Bottleneck via Statistical-Manifold Compression
Weiqi Wang, Zhiyi Tian, Chenhan Zhang +1
Information Bottleneck (IB) is widely used, but in deep learning, it is usually implemented through tractable surrogates, such as variational bounds or neural mutual information (M…
EVE: Efficient Verification of Data Erasure through Customized Perturbation in Approximate Unlearning
Weiqi Wang, Zhiyi Tian, Chenhan Zhang +2
Verifying whether the machine unlearning process has been properly executed is critical but remains underexplored. Some existing approaches propose unlearning verification methods…
BlindU: Blind Machine Unlearning without Revealing Erasing Data
Weiqi Wang, Zhiyi Tian, Chenhan Zhang +1
Machine unlearning enables data holders to remove the contribution of their specified samples from trained models to protect their privacy. However, it is paradoxical that most unl…