Backdoor Attacks and Countermeasures on Deep Learning: A Comprehensive Review
arXiv:2007.10760
Abstract
This work provides the community with a timely comprehensive review of backdoor attacks and countermeasures on deep learning. According to the attacker's capability and affected stage of the machine learning pipeline, the attack surfaces are recognized to be wide and then formalized into six categorizations: code poisoning, outsourcing, pretrained, data collection, collaborative learning and post-deployment. Accordingly, attacks under each categorization are combed. The countermeasures are categorized into four general classes: blind backdoor removal, offline backdoor inspection, online backdoor inspection, and post backdoor removal. Accordingly, we review countermeasures, and compare and analyze their advantages and disadvantages. We have also reviewed the flip side of backdoor attacks, which are explored for i) protecting intellectual property of deep learning models, ii) acting as a honeypot to catch adversarial example attacks, and iii) verifying data deletion requested by the data contributor.Overall, the research on defense is far behind the attack, and there is no single defense that can prevent all types of backdoor attacks. In some cases, an attacker can intelligently bypass existing defenses with an adaptive attack. Drawing the insights from the systematic review, we also present key areas for future research on the backdoor, such as empirical security evaluations from physical trigger attacks, and in particular, more efficient and practical countermeasures are solicited.
29 pages, 9 figures, 2 tables
References in corpus (22)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning
- Towards Federated Learning at Scale: System Design
- Can You Really Backdoor Federated Learning?
- A Berkeley View of Systems Challenges for AI
- Transferable Clean-Label Poisoning Attacks on Deep Neural Nets
- Robust Anomaly Detection and Backdoor Attack Detection Via Differential Privacy
- NeuronInspect: Detecting Backdoors in Neural Networks via Output Explanations
- Defending Neural Backdoors via Generative Distribution Modeling
- Weight Poisoning Attacks on Pre-trained Models
- Label-Consistent Backdoor Attacks
- BlackMarks: Blackbox Multibit Watermarking for Deep Neural Networks
- Backdoor attacks and defenses in feature-partitioned collaborative learning
- ConFoc: Content-Focus Protection Against Trojan Attacks on Neural Networks
- Design of intentional backdoors in sequential models
- Poison as a Cure: Detecting & Neutralizing Variable-Sized Backdoor Attacks in Deep Neural Networks
- Programmable Neural Network Trojan for Pre-Trained Feature Extractor
- Design and Evaluation of a Multi-Domain Trojan Detection Method on Deep Neural Networks
- Adversarial Audio: A New Information Hiding Method and Backdoor for DNN-based Speech Recognition Models
- FaceHack: Triggering backdoored facial recognition systems using facial characteristics
- Live Trojan Attacks on Deep Neural Networks
- AI Data poisoning attack: Manipulating game AI of Go