4 citations · 8 across the 13 of their papers we have counts for
7 papers · 1 filter
Steering Beyond the Support: Adversarial Training on Unsupervised Jailbroken Activation Simulation
Luoyu Chen, Weiqi Wang, Zhiyi Tian +5
Jailbreak prompts can trigger harmful completions on aligned LLMs, In accordance, safety steering has been proposed: test-time activation interventions that steer jailbreak activat…
Ellipsoid Control: A White-list Jailbreak Defense via Benign Latent Modeling
Luoyu Chen, Weiqi Wang, Zhiyi Tian +3
Representation engineering (RepE) defenses have shown strong robustness against jailbreak attacks on large language models (LLMs). However, these methods fundamentally rely on blac…
BlindU: Blind Machine Unlearning without Revealing Erasing Data
Weiqi Wang, Zhiyi Tian, Chenhan Zhang +1
Machine unlearning enables data holders to remove the contribution of their specified samples from trained models to protect their privacy. However, it is paradoxical that most unl…
TAPE: Tailored Posterior Difference for Auditing of Machine Unlearning
Weiqi Wang, Zhiyi Tian, An Liu +1
With the increasing prevalence of Web-based platforms handling vast amounts of user data, machine unlearning has emerged as a crucial mechanism to uphold users' right to be forgott…
CRFU: Compressive Representation Forgetting Against Privacy Leakage on Machine Unlearning
Weiqi Wang, Chenhan Zhang, Zhiyi Tian +2
Machine unlearning allows data owners to erase the impact of their specified data from trained models. Unfortunately, recent studies have shown that adversaries can recover the era…
SCU: An Efficient Machine Unlearning Scheme for Deep Learning Enabled Semantic Communications
Weiqi Wang, Zhiyi Tian, Chenhan Zhang +1
Deep learning (DL) enabled semantic communications leverage DL to train encoders and decoders (codecs) to extract and recover semantic information. However, most semantic training…