11 papers
Steering Beyond the Support: Adversarial Training on Unsupervised Jailbroken Activation Simulation
Luoyu Chen, Weiqi Wang, Zhiyi Tian +5
Jailbreak prompts can trigger harmful completions on aligned LLMs, In accordance, safety steering has been proposed: test-time activation interventions that steer jailbreak activat…
Ellipsoid Control: A White-list Jailbreak Defense via Benign Latent Modeling
Luoyu Chen, Weiqi Wang, Zhiyi Tian +3
Representation engineering (RepE) defenses have shown strong robustness against jailbreak attacks on large language models (LLMs). However, these methods fundamentally rely on blac…
Approximate Machine Unlearning through Manifold Representation Forgetting Guided by Self Mode Connectivity
Weiqi Wang, Zhiyi Tian, Chenhan Zhang +2
Machine unlearning is a fundamental mechanism that enforces the right to be forgotten. Existing unlearning studies that rely on label manipulation or task-gradient reversal often d…
Machine Unlearning: A Comprehensive Survey
Weiqi Wang, Zhiyi Tian, Chenhan Zhang +1
As the right to be forgotten has been legislated worldwide, many studies attempt to design unlearning mechanisms to protect users' privacy when they want to leave machine learning…
GeoIB: Geometry-Aware Information Bottleneck via Statistical-Manifold Compression
Weiqi Wang, Zhiyi Tian, Chenhan Zhang +1
Information Bottleneck (IB) is widely used, but in deep learning, it is usually implemented through tractable surrogates, such as variational bounds or neural mutual information (M…
EVE: Efficient Verification of Data Erasure through Customized Perturbation in Approximate Unlearning
Weiqi Wang, Zhiyi Tian, Chenhan Zhang +2
Verifying whether the machine unlearning process has been properly executed is critical but remains underexplored. Some existing approaches propose unlearning verification methods…