Defending against Backdoor Attack on Deep Neural Networks
arXiv:2002.12162
Abstract
Although deep neural networks (DNNs) have achieved a great success in various computer vision tasks, it is recently found that they are vulnerable to adversarial attacks. In this paper, we focus on the so-called \textit{backdoor attack}, which injects a backdoor trigger to a small portion of training data (also known as data poisoning) such that the trained DNN induces misclassification while facing examples with this trigger. To be specific, we carefully study the effect of both real and synthetic backdoor attacks on the internal response of vanilla and backdoored DNNs through the lens of Gard-CAM. Moreover, we show that the backdoor attack induces a significant bias in neuron activation in terms of the norm of an activation map compared to its and norm. Spurred by our results, we propose the \textit{-based neuron pruning} to remove the backdoor from the backdoored DNN. Experiments show that our method could effectively decrease the attack success rate, and also hold a high classification accuracy for clean images.
This workshop manuscript is not a publication and will not be published anywhere
References in corpus (5)
- BadNets: Identifying Vulnerabilities in the Machine Learning Model Supply Chain
- Spectral Signatures in Backdoor Attacks
- Is feature selection secure against training data poisoning?
- Structured Adversarial Attack: Towards General Implementation and Better Interpretability
- Interpreting Adversarial Examples by Activation Promotion and Suppression
Cited by in corpus (7)
- Input-Aware Dynamic Backdoor Attack
- Wild Patterns Reloaded: A Survey of Machine Learning Security against Training Data Poisoning
- WaNet -- Imperceptible Warping-based Backdoor Attack
- Natural Backdoor Attack on Text Data
- Practical Detection of Trojan Neural Networks: Data-Limited and Data-Free Cases
- Light Can Hack Your Face! Black-box Backdoor Attack on Face Recognition Systems
- FIBA: Frequency-Injection based Backdoor Attack in Medical Image Analysis