Mitigating Advanced Adversarial Attacks with More Advanced Gradient Obfuscation Techniques
arXiv:2005.13712
Abstract
Deep Neural Networks (DNNs) are well-known to be vulnerable to Adversarial Examples (AEs). A large amount of efforts have been spent to launch and heat the arms race between the attackers and defenders. Recently, advanced gradient-based attack techniques were proposed (e.g., BPDA and EOT), which have defeated a considerable number of existing defense methods. Up to today, there are still no satisfactory solutions that can effectively and efficiently defend against those attacks. In this paper, we make a steady step towards mitigating those advanced gradient-based attacks with two major contributions. First, we perform an in-depth analysis about the root causes of those attacks, and propose four properties that can break the fundamental assumptions of those attacks. Second, we identify a set of operations that can meet those properties. By integrating these operations, we design two preprocessing functions that can invalidate these powerful attacks. Extensive evaluations indicate that our solutions can effectively mitigate all existing standard and advanced attack techniques, and beat 11 state-of-the-art defense solutions published in top-tier conferences over the past 2 years. The defender can employ our solutions to constrain the attack success rate below 7% for the strongest attacks even the adversary has spent dozens of GPU hours.
References in corpus (12)
- Feature Squeezing: Detecting Adversarial Examples in Deep Neural Networks
- Towards Deep Neural Network Architectures Robust to Adversarial Examples
- On Evaluating Adversarial Robustness
- Technical Report on the CleverHans v2.1.0 Adversarial Examples Library
- Understanding Adversarial Training: Increasing Local Stability of Neural Nets through Robust Optimization
- PixelDefend: Leveraging Generative Models to Understand and Defend against Adversarial Examples
- Improving the Adversarial Robustness and Interpretability of Deep Neural Networks by Regularizing their Input Gradients
- Learning with a Strong Adversary
- MagNet and "Efficient Defenses Against Adversarial Attacks" are Not Robust to Adversarial Examples
- On Adaptive Attacks to Adversarial Example Defenses
- Generative Adversarial Trainer: Defense to Adversarial Perturbations with GAN
- Boosting Adversarial Attacks with Momentum
Cited by in corpus (5)
- DeepSweep: An Evaluation Framework for Mitigating DNN Backdoor Attacks using Data Augmentation
- FenceBox: A Platform for Defeating Adversarial Examples with Data Augmentation Techniques
- The Feasibility and Inevitability of Stealth Attacks
- A Data Augmentation-based Defense Method Against Adversarial Attacks in Neural Networks
- Privacy-preserving Collaborative Learning with Automatic Transformation Search