#machine unlearning
5 papers match
Benign on Label, Malicious by Design: Clean-Label Dormant-to-Activated Backdoor via Machine Unlearning with Removable Camouflage
Dongdong Zhao, Can Li, Xiang Yao +3
The paper proposes a clean‑label backdoor attack that stays dormant during training and becomes active only after specific camouflage samples are removed via machine unlearning, us…
Beyond Binary Rewards: A Comparative Study of Reward Design for Reinforcement Unlearning
Efstratios Zaradoukas, Davide Gabrielli, Bardh Prenkaj +1
The paper investigates how different reward functions affect the speed and effectiveness of reinforcement‑learning based machine unlearning for language models, proposing graded an…
When Machine Unlearning Meets Retrieval-Augmented Generation (RAG): Keep Secret or Forget Knowledge?
Shang Wang, Tianqing Zhu, Dayong Ye +1
The paper proposes a lightweight method to make large language models forget specific information by altering the external knowledge base of Retrieval‑Augmented Generation systems,…
Inference-Time Machine Unlearning via Gated Activation Redirection
VinÃcius Conte Turani, Otávio Parraga, João Vitor Boer Abitante +7
The paper proposes GUARD-IT, a gradient‑free method that modifies activations at inference time with input‑dependent rotations to erase specific data from large language models whi…
Signal-Guided Optimization for Machine Unlearning
Xujia Li, Dan Li, Jian Lou +1
The paper introduces GSUO, a guidance-signal-aware optimization framework that uses fine-grained task-specific signals to improve the effectiveness and efficiency of machine unlear…