Two Sides of the Same Coin: Learning the Backdoor to Remove the Backdoor
arXiv:2607.05748 · doi:10.1609/aaai.v39i21.34441
Abstract
The community has recently developed various training-time defenses to counter neural backdoors introduced through data poisoning. In light of the observation that a model learns poisonous samples responsible for the backdoor easier than benign samples, these approaches either use a fixed threshold of the training loss for splitting or iteratively learn a reference model as an oracle for identifying benign samples. In particular, the latter has proven effective for anti-backdoor learning. Our method, HARVEY, leverages a similar yet crucially different technique: learning an oracle for poisonous rather than benign samples. Learning a backdoored reference model is significantly easier than learning a reference model on benign data. Consequently, we can identify poisonous samples much more accurately than related work identifies benign samples. This crucial difference enables near-perfect backdoor removal as we demonstrate in our evaluation. HARVEY substantially outperforms related approaches across attack types, datasets, and architectures, lowering the attack success rate to the very minimum at a negligible loss in natural accuracy. The figure below shows an overview of our methods working principle.
References in corpus (13)
- Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning
- Neural Attention Distillation: Erasing Backdoor Triggers from Deep Neural Networks
- BackdoorBench: A Comprehensive Benchmark of Backdoor Learning
- Backdoor Defense via Decoupling the Training Process
- Anti-Backdoor Learning: Training Clean Models on Poisoned Data
- Bridging Mode Connectivity in Loss Landscapes and Adversarial Robustness
- DeepSweep: An Evaluation Framework for Mitigating DNN Backdoor Attacks using Data Augmentation
- Practical Detection of Trojan Neural Networks: Data-Limited and Data-Free Cases
- SPECTRE: Defending Against Backdoor Attacks Using Robust Statistics
- UNICORN: A Unified Backdoor Trigger Inversion Framework
- Backdoor Defense via Adaptively Splitting Poisoned Dataset
- DataElixir: Purifying Poisoned Dataset to Mitigate Backdoor Attacks via Diffusion Models
- Backdoor Defense via Deconfounded Representation Learning