Reflection Backdoor: A Natural Backdoor Attack on Deep Neural Networks
arXiv:2007.02343
Abstract
Recent studies have shown that DNNs can be compromised by backdoor attacks crafted at training time. A backdoor attack installs a backdoor into the victim model by injecting a backdoor pattern into a small proportion of the training data. At test time, the victim model behaves normally on clean test data, yet consistently predicts a specific (likely incorrect) target class whenever the backdoor pattern is present in a test example. While existing backdoor attacks are effective, they are not stealthy. The modifications made on training data or labels are often suspicious and can be easily detected by simple data filtering or human inspection. In this paper, we present a new type of backdoor attack inspired by an important natural phenomenon: reflection. Using mathematical modeling of physical reflection models, we propose reflection backdoor (Refool) to plant reflections as backdoor into a victim model. We demonstrate on 3 computer vision tasks and 5 datasets that, Refool can attack state-of-the-art DNNs with high success rate, and is resistant to state-of-the-art backdoor defenses.
Accepted by ECCV-2020
References in corpus (18)
- Sequence to Sequence Learning with Neural Networks
- mixup: Beyond Empirical Risk Minimization
- Understanding Black-box Predictions via Influence Functions
- Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning
- BadNets: Identifying Vulnerabilities in the Machine Learning Model Supply Chain
- Understanding Adversarial Attacks on Deep Learning Based Medical Image Analysis Systems
- Robust Physical-World Attacks on Deep Learning Models
- Countering Adversarial Images using Input Transformations
- Can You Really Backdoor Federated Learning?
- Februus: Input Purification Defense Against Trojan Attacks on Deep Neural Network Systems
- Detecting Backdoor Attacks on Deep Neural Networks by Activation Clustering
- Support Vector Machines under Adversarial Label Contamination
- Natural Adversarial Examples
- On the Convergence and Robustness of Adversarial Training
- Skip Connections Matter: On the Transferability of Adversarial Examples Generated with ResNets
- Backdoor Embedding in Convolutional Neural Network Models via Invisible Perturbation
- Rethinking the Trigger of Backdoor Attack
- Semantic Guided Single Image Reflection Removal
Cited by in corpus (5)
- WaNet -- Imperceptible Warping-based Backdoor Attack
- Poisoned classifiers are not only backdoored, they are fundamentally broken
- Sneaky Spikes: Uncovering Stealthy Backdoor Attacks in Spiking Neural Networks with Neuromorphic Data
- LMSanitator: Defending Prompt-Tuning Against Task-Agnostic Backdoors
- Gradient Shaping: Enhancing Backdoor Attack Against Reverse Engineering