PixelDefend: Leveraging Generative Models to Understand and Defend against Adversarial Examples
arXiv:1710.10766
Abstract
Adversarial perturbations of normal images are usually imperceptible to humans, but they can seriously confuse state-of-the-art machine learning models. What makes them so special in the eyes of image classifiers? In this paper, we show empirically that adversarial examples mainly lie in the low probability regions of the training distribution, regardless of attack types and targeted models. Using statistical hypothesis testing, we find that modern neural density models are surprisingly good at detecting imperceptible image perturbations. Based on this discovery, we devised PixelDefend, a new approach that purifies a maliciously perturbed image by moving it back towards the distribution seen in the training data. The purified image is then run through an unmodified classifier, making our method agnostic to both the classifier and the attacking method. As a result, PixelDefend can be used to protect already deployed models and be combined with other model-specific defenses. Experiments show that our method greatly improves resilience across a wide variety of state-of-the-art attacking methods, increasing accuracy on the strongest attack from 63% to 84% for Fashion MNIST and from 32% to 70% for CIFAR-10.
ICLR 2018
References in corpus (7)
- Fashion-MNIST: a Novel Image Dataset for Benchmarking Machine Learning Algorithms
- Delving into Transferable Adversarial Examples and Black-box Attacks
- Towards Deep Neural Network Architectures Robust to Adversarial Examples
- The Space of Transferable Adversarial Examples
- Biologically inspired protection of deep networks from adversarial attacks
- Feature Squeezing Mitigates and Detects Carlini/Wagner Adversarial Examples
- Clustering via Mode Seeking by Direct Estimation of the Gradient of a Log-Density
Cited by in corpus (110)
- Obfuscated Gradients Give a False Sense of Security: Circumventing Defenses to Adversarial Examples
- Fast is better than free: Revisiting adversarial training
- Adversarial Risk and the Dangers of Evaluating Against Weak Attacks
- Adversarial Examples: Attacks and Defenses for Deep Learning
- Are Labels Required for Improving Adversarial Robustness?
- Threat of Adversarial Attacks on Deep Learning in Computer Vision: A Survey
- Adversarial Machine Learning in Image Classification: A Survey Towards the Defender's Perspective
- Motivating the Rules of the Game for Adversarial Example Research
- Adversarial Sample Detection for Deep Neural Network through Model Mutation Testing
- Adversarial Attack Vulnerability of Medical Image Analysis Systems: Unexplored Factors
- Defense Against Adversarial Attacks Using Feature Scattering-based Adversarial Training
- The Robust Manifold Defense: Adversarial Training using Generative Models
- A Review of Adversarial Attack and Defense for Classification Methods
- Your Classifier is Secretly an Energy Based Model and You Should Treat it Like One
- Improving Transferability of Adversarial Examples with Input Diversity
- You Only Propagate Once: Accelerating Adversarial Training via Maximal Principle
- Test-Time Training with Self-Supervision for Generalization under Distribution Shifts
- The Limitations of Adversarial Training and the Blind-Spot Attack
- Security and Privacy Issues in Deep Learning
- Convergence of Adversarial Training in Overparametrized Neural Networks
- Overfitting in adversarially robust deep learning
- Adversarial Examples - A Complete Characterisation of the Phenomenon
- Towards Robust Neural Networks via Random Self-ensemble
- Towards Stable and Efficient Training of Verifiably Robust Neural Networks
- Adversarially Robust Neural Architectures
- Do Wider Neural Networks Really Help Adversarial Robustness?
- Daedalus: Breaking Non-Maximum Suppression in Object Detection via Adversarial Examples
- Defensive Quantization: When Efficiency Meets Robustness
- DeepHunter: Hunting Deep Neural Network Defects via Coverage-Guided Fuzzing
- Robust Neural Networks using Randomized Adversarial Training
- The Threat of Adversarial Attacks on Machine Learning in Network Security -- A Survey
- Feature Distillation: DNN-Oriented JPEG Compression Against Adversarial Examples
- Evolving Robust Neural Architectures to Defend from Adversarial Attacks
- Regional Homogeneity: Towards Learning Transferable Universal Adversarial Perturbations Against Defenses
- ComDefend: An Efficient Image Compression Model to Defend Adversarial Examples
- DefenseVGAE: Defending against Adversarial Attacks on Graph Data via a Variational Graph Autoencoder
- Unsupervised Out-of-Distribution Detection with Batch Normalization
- Mitigating Advanced Adversarial Attacks with More Advanced Gradient Obfuscation Techniques
- Breaking Transferability of Adversarial Samples with Randomness
- Robustifying Models Against Adversarial Attacks by Langevin Dynamics
- Investigating Robustness of Adversarial Samples Detection for Automatic Speaker Verification
- Bridging the Performance Gap between FGSM and PGD Adversarial Training
- Enhancing Gradient-based Attacks with Symbolic Intervals
- Adversarial Examples on Object Recognition: A Comprehensive Survey
- Detecting Adversarial Examples in Convolutional Neural Networks
- Distributionally Adversarial Attack
- Retrieval-Augmented Convolutional Neural Networks for Improved Robustness against Adversarial Examples
- On Certifying Non-uniform Bound against Adversarial Attacks
- Input Validation for Neural Networks via Runtime Local Robustness Verification
- Adaptive Adversarial Attack on Scene Text Recognition
- Using Undervolting as an On-Device Defense Against Adversarial Machine Learning Attacks
- A Person Re-identification Data Augmentation Method with Adversarial Defense Effect
- Better the Devil you Know: An Analysis of Evasion Attacks using Out-of-Distribution Adversarial Examples
- An Effective and Robust Detector for Logo Detection
- APRICOT: A Dataset of Physical Adversarial Attacks on Object Detection
- CAGFuzz: Coverage-Guided Adversarial Generative Fuzzing Testing of Deep Learning Systems
- Attacking and Defending Machine Learning Applications of Public Cloud
- Improving Global Adversarial Robustness Generalization With Adversarially Trained GAN
- CAP-GAN: Towards Adversarial Robustness with Cycle-consistent Attentional Purification
- On the Need for Topology-Aware Generative Models for Manifold-Based Defenses
- Effects of Loss Functions And Target Representations on Adversarial Robustness
- Strength in Numbers: Trading-off Robustness and Computation via Adversarially-Trained Ensembles
- Feature Prioritization and Regularization Improve Standard Accuracy and Adversarial Robustness
- Defending against Adversarial Attack towards Deep Neural Networks via Collaborative Multi-task Training
- Detecting Localized Adversarial Examples: A Generic Approach using Critical Region Analysis
- Towards Robust Neural Networks via Close-loop Control
- Adversarial Defense by Stratified Convolutional Sparse Coding
- Predicting Adversarial Examples with High Confidence
- Towards Understanding Limitations of Pixel Discretization Against Adversarial Attacks
- A Game Theoretic Analysis of Additive Adversarial Attacks and Defenses
- Less is More: Robust and Novel Features for Malicious Domain Detection
- Understanding Catastrophic Overfitting in Adversarial Training
- Improving Model Robustness with Latent Distribution Locally and Globally
- Designing Adversarially Resilient Classifiers using Resilient Feature Engineering
- Drawing Robust Scratch Tickets: Subnetworks with Inborn Robustness Are Found within Randomly Initialized Networks
- Auditing AI models for Verified Deployment under Semantic Specifications
- RAIN: A Simple Approach for Robust and Accurate Image Classification Networks
- Self-supervised Adversarial Training
- Purifying Adversarial Perturbation with Adversarially Trained Auto-encoders
- DAFAR: Defending against Adversaries by Feedback-Autoencoder Reconstruction
- Defective Convolutional Networks
- Applying Tensor Decomposition to image for Robustness against Adversarial Attack
- Attack as Defense: Characterizing Adversarial Examples using Robustness
- AdvFoolGen: Creating Persistent Troubles for Deep Classifiers
- Implicit Generative Modeling of Random Noise during Training for Adversarial Robustness
- Distributional Robustness with IPMs and links to Regularization and GANs
- Efficient detection of adversarial images
- The Vulnerability of the Neural Networks Against Adversarial Examples in Deep Learning Algorithms
- Local Competition and Stochasticity for Adversarial Robustness in Deep Learning
- The art of defense: letting networks fool the attacker
- When the Guard failed the Droid: A case study of Android malware
- Biologically Inspired Visual System Architecture for Object Recognition in Autonomous Systems
- Local Competition and Uncertainty for Adversarial Robustness in Deep Learning
- A Survey on Assessing the Generalization Envelope of Deep Neural Networks: Predictive Uncertainty, Out-of-distribution and Adversarial Samples
- A Useful Taxonomy for Adversarial Robustness of Neural Networks
- DOI: Divergence-based Out-of-Distribution Indicators via Deep Generative Models
- Generative Models for Security: Attacks, Defenses, and Opportunities
- Reject Illegal Inputs with Generative Classifier Derived from Any Discriminative Classifier
- Detecting Adversaries, yet Faltering to Noise? Leveraging Conditional Variational AutoEncoders for Adversary Detection in the Presence of Noisy Images
- Shape Defense Against Adversarial Attacks
- MadNet: Using a MAD Optimization for Defending Against Adversarial Attacks
- An Efficient Pre-processing Method to Eliminate Adversarial Effects
- Testing for Typicality with Respect to an Ensemble of Learned Distributions
- Adversarial Data Encryption
- Defending Against Adversarial Attacks Using Random Forests
- Defense Through Diverse Directions
- Beyond Categorical Label Representations for Image Classification
- On Intrinsic Dataset Properties for Adversarial Machine Learning
- Harnessing adversarial examples with a surprisingly simple defense
- Enhancing Transformation-based Defenses using a Distribution Classifier