A General Framework for Defending Against Backdoor Attacks via Influence Graph
arXiv:2111.14309
Abstract
In this work, we propose a new and general framework to defend against backdoor attacks, inspired by the fact that attack triggers usually follow a \textsc{specific} type of attacking pattern, and therefore, poisoned training examples have greater impacts on each other during training. We introduce the notion of the {\it influence graph}, which consists of nodes and edges respectively representative of individual training points and associated pair-wise influences. The influence between a pair of training points represents the impact of removing one training point on the prediction of another, approximated by the influence function \citep{koh2017understanding}. Malicious training points are extracted by finding the maximum average sub-graph subject to a particular size. Extensive experiments on computer vision and natural language processing tasks demonstrate the effectiveness and generality of the proposed framework.
References in corpus (15)
- Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning
- Poison Ink: Robust and Invisible Backdoor Attack
- On the Effectiveness of Mitigating Data Poisoning Attacks with Gradient Shaping
- Robust Anomaly Detection and Backdoor Attack Detection Via Differential Privacy
- Weight Poisoning Attacks on Pre-trained Models
- Label-Consistent Backdoor Attacks
- BadPre: Task-agnostic Backdoor Attacks to Pre-trained NLP Foundation Models
- ConFoc: Content-Focus Protection Against Trojan Attacks on Neural Networks
- Hidden Killer: Invisible Textual Backdoor Attacks with Syntactic Trigger
- SPECTRE: Defending Against Backdoor Attacks Using Robust Statistics
- Be Careful about Poisoned Word Embeddings: Exploring the Vulnerability of the Embedding Layers in NLP Models
- Defending Against Backdoor Attacks in Natural Language Generation
- Pair the Dots: Jointly Examining Training History and Test Stimuli for Model Interpretability
- Turn the Combination Lock: Learnable Textual Backdoor Attacks via Word Substitution
- Spinning Sequence-to-Sequence Models with Meta-Backdoors