Gradient Shaping: Enhancing Backdoor Attack Against Reverse Engineering
arXiv:2301.12318 · doi:10.14722/ndss.2024.24450
Abstract
Most existing methods to detect backdoored machine learning (ML) models take one of the two approaches: trigger inversion (aka. reverse engineer) and weight analysis (aka. model diagnosis). In particular, the gradient-based trigger inversion is considered to be among the most effective backdoor detection techniques, as evidenced by the TrojAI competition, Trojan Detection Challenge and backdoorBench. However, little has been done to understand why this technique works so well and, more importantly, whether it raises the bar to the backdoor attack. In this paper, we report the first attempt to answer this question by analyzing the change rate of the backdoored model around its trigger-carrying inputs. Our study shows that existing attacks tend to inject the backdoor characterized by a low change rate around trigger-carrying inputs, which are easy to capture by gradient-based trigger inversion. In the meantime, we found that the low change rate is not necessary for a backdoor attack to succeed: we design a new attack enhancement called \textit{Gradient Shaping} (GRASP), which follows the opposite direction of adversarial training to reduce the change rate of a backdoored model with regard to the trigger, without undermining its backdoor effect. Also, we provide a theoretic analysis to explain the effectiveness of this new technique and the fundamental weakness of gradient-based trigger inversion. Finally, we perform both theoretical and experimental analysis, showing that the GRASP enhancement does not reduce the effectiveness of the stealthy attacks against the backdoor detection methods based on weight analysis, as well as other backdoor mitigation methods without using detection.
References in corpus (24)
- Adam: A Method for Stochastic Optimization
- Intriguing properties of neural networks
- BadNets: Identifying Vulnerabilities in the Machine Learning Model Supply Chain
- Detecting Backdoor Attacks on Deep Neural Networks by Activation Clustering
- Neural Attention Distillation: Erasing Backdoor Triggers from Deep Neural Networks
- TABOR: A Highly Accurate Approach to Inspecting and Restoring Trojan Backdoors in AI Systems
- Benchmarking Robustness in Object Detection: Autonomous Driving when Winter is Coming
- WaNet -- Imperceptible Warping-based Backdoor Attack
- Blind Backdoors in Deep Learning Models
- Training Quantized Nets: A Deeper Understanding
- Robust Anomaly Detection and Backdoor Attack Detection Via Differential Privacy
- Backdoor Scanning for Deep Neural Networks through K-Arm Optimization
- RAB: Provable Robustness Against Backdoor Attacks
- Backdoor Defense via Decoupling the Training Process
- Reflection Backdoor: A Natural Backdoor Attack on Deep Neural Networks
- Anti-Backdoor Learning: Training Clean Models on Poisoned Data
- DeepSweep: An Evaluation Framework for Mitigating DNN Backdoor Attacks using Data Augmentation
- The TrojAI Software Framework: An OpenSource tool for Embedding Trojans into Deep Learning Models
- Adversarial Lipschitz Regularization
- HaS-Nets: A Heal and Select Mechanism to Defend DNNs Against Backdoor Attacks for Data Collection Scenarios
- Circumventing Backdoor Defenses That Are Based on Latent Separability
- Lipschitz Bounds and Provably Robust Training by Laplacian Smoothing
- Under-confidence Backdoors Are Resilient and Stealthy Backdoors
- Trojan Horse Training for Breaking Defenses against Backdoor Attacks in Deep Learning