MagNet: a Two-Pronged Defense against Adversarial Examples
arXiv:1705.09064
Abstract
Deep learning has shown promising results on hard perceptual problems in recent years. However, deep learning systems are found to be vulnerable to small adversarial perturbations that are nearly imperceptible to human. Such specially crafted perturbations cause deep learning systems to output incorrect decisions, with potentially disastrous consequences. These vulnerabilities hinder the deployment of deep learning systems where safety or security is important. Attempts to secure deep learning systems either target specific attacks or have been shown to be ineffective. In this paper, we propose MagNet, a framework for defending neural network classifiers against adversarial examples. MagNet does not modify the protected classifier or know the process for generating adversarial examples. MagNet includes one or more separate detector networks and a reformer network. Different from previous work, MagNet learns to differentiate between normal and adversarial examples by approximating the manifold of normal examples. Since it does not rely on any process for generating adversarial examples, it has substantial generalization power. Moreover, MagNet reconstructs adversarial examples by moving them towards the manifold, which is effective for helping classify adversarial examples with small perturbation correctly. We discuss the intrinsic difficulty in defending against whitebox attack and propose a mechanism to defend against graybox attack. Inspired by the use of randomness in cryptography, we propose to use diversity to strengthen MagNet. We show empirically that MagNet is effective against most advanced state-of-the-art attacks in blackbox and graybox scenarios while keeping false positive rate on normal examples very low.
Accepted at the ACM Conference on Computer and Communications Security (CCS), 2017
References in corpus (8)
- Striving for Simplicity: The All Convolutional Net
- Delving into Transferable Adversarial Examples and Black-box Attacks
- Uncertainty-Aware Reinforcement Learning for Collision Avoidance
- Adversarial Examples for Semantic Segmentation and Object Detection
- Universal adversarial perturbations
- Adversarial examples for generative models
- Learning Transferable Policies for Monocular Reactive MAV Control
- Deep Visual Foresight for Planning Robot Motion
Cited by in corpus (76)
- Feature Squeezing: Detecting Adversarial Examples in Deep Neural Networks
- Wild Patterns: Ten Years After the Rise of Adversarial Machine Learning
- Obfuscated Gradients Give a False Sense of Security: Circumventing Defenses to Adversarial Examples
- Adversarial Risk and the Dangers of Evaluating Against Weak Attacks
- Deep k-Nearest Neighbors: Towards Confident, Interpretable and Robust Deep Learning
- Februus: Input Purification Defense Against Trojan Attacks on Deep Neural Network Systems
- Adversarial Examples: Opportunities and Challenges
- Invisible for both Camera and LiDAR: Security of Multi-Sensor Fusion based Perception in Autonomous Driving Under Physical-World Attacks
- Mitigating Adversarial Effects Through Randomization
- Adversarial Example Detection for DNN Models: A Review and Experimental Comparison
- Threat of Adversarial Attacks on Deep Learning in Computer Vision: A Survey
- Evaluating the Robustness of Neural Networks: An Extreme Value Theory Approach
- MagNet and "Efficient Defenses Against Adversarial Attacks" are Not Robust to Adversarial Examples
- The Robust Manifold Defense: Adversarial Training using Generative Models
- Gotta Catch 'Em All: Using Honeypots to Catch Adversarial Attacks on Neural Networks
- Robust Pre-Training by Adversarial Contrastive Learning
- Malware Makeover: Breaking ML-based Static Analysis by Modifying Executable Bytes
- Detecting and Diagnosing Adversarial Images with Class-Conditional Capsule Reconstructions
- Detecting Adversarial Examples by Input Transformations, Defense Perturbations, and Voting
- Towards Robust Neural Networks via Random Self-ensemble
- Detecting Adversarial Attacks on Neural Network Policies with Visual Foresight
- Daedalus: Breaking Non-Maximum Suppression in Object Detection via Adversarial Examples
- Adversarial Neural Network Inversion via Auxiliary Knowledge Alignment
- Predify: Augmenting deep neural networks with brain-inspired predictive coding dynamics
- Defending Against Adversarial Attacks by Leveraging an Entire GAN
- Defend Data Poisoning Attacks on Voice Authentication
- The Faults in Our Pi Stars: Security Issues and Open Challenges in Deep Reinforcement Learning
- Defense against Adversarial Attacks Using High-Level Representation Guided Denoiser
- What You See is Not What the Network Infers: Detecting Adversarial Examples Based on Semantic Contradiction
- PAC-learning in the presence of evasion adversaries
- Interpretable Deep Learning under Fire
- Adversarial Defense by Latent Style Transformations
- InfoAT: Improving Adversarial Training Using the Information Bottleneck Principle
- VerIDeep: Verifying Integrity of Deep Neural Networks through Sensitive-Sample Fingerprinting
- GreedyFool: Distortion-Aware Sparse Adversarial Attack
- Interpretability is a Kind of Safety: An Interpreter-based Ensemble for Adversary Defense
- Featurized Bidirectional GAN: Adversarial Defense via Adversarially Learned Semantic Inference
- Post-breach Recovery: Protection against White-box Adversarial Examples for Leaked DNN Models
- Fooling Vision and Language Models Despite Localization and Attention Mechanism
- Exploiting Vulnerabilities of Deep Learning-based Energy Theft Detection in AMI through Adversarial Attacks
- Input Validation for Neural Networks via Runtime Local Robustness Verification
- Detecting Adversarial Examples via Key-based Network
- SUDS: Sanitizing Universal and Dependent Steganography
- Better the Devil you Know: An Analysis of Evasion Attacks using Out-of-Distribution Adversarial Examples
- Deep Learning in Information Security
- Gradient Similarity: An Explainable Approach to Detect Adversarial Attacks against Deep Learning
- A Data-driven Adversarial Examples Recognition Framework via Adversarial Feature Genome
- Defending against Adversarial Attack towards Deep Neural Networks via Collaborative Multi-task Training
- The Adversarial Attack and Detection under the Fisher Information Metric
- Optimal Transport Classifier: Defending Against Adversarial Attacks by Regularized Deep Embedding
- Machine vs Machine: Minimax-Optimal Defense Against Adversarial Examples
- Justification-Based Reliability in Machine Learning
- Stochastic Combinatorial Ensembles for Defending Against Adversarial Examples
- Stealing Deep Reinforcement Learning Models for Fun and Profit
- The Dimpled Manifold Model of Adversarial Examples in Machine Learning
- AutoGAN: Robust Classifier Against Adversarial Attacks
- A Mixture Model Based Defense for Data Poisoning Attacks Against Naive Bayes Spam Filters
- Stroke-based Character Reconstruction
- FineFool: Fine Object Contour Attack via Attention
- AdvCodeMix: Adversarial Attack on Code-Mixed Data
- Iterative Window Mean Filter: Thwarting Diffusion-based Adversarial Purification
- Automated Detection System for Adversarial Examples with High-Frequency Noises Sieve
- Less is More: Culling the Training Set to Improve Robustness of Deep Neural Networks
- ZK-GanDef: A GAN based Zero Knowledge Adversarial Training Defense for Neural Networks
- Classification Auto-Encoder based Detector against Diverse Data Poisoning Attacks
- A Hierarchical Feature Constraint to Camouflage Medical Adversarial Attacks
- On The Utility of Conditional Generation Based Mutual Information for Characterizing Adversarial Subspaces
- MixDefense: A Defense-in-Depth Framework for Adversarial Example Detection Based on Statistical and Semantic Analysis
- Unifying Bilateral Filtering and Adversarial Training for Robust Neural Networks
- FBI: Fingerprinting models with Benign Inputs
- Search Space of Adversarial Perturbations against Image Filters
- ManiGen: A Manifold Aided Black-box Generator of Adversarial Examples
- Detecting Adversaries, yet Faltering to Noise? Leveraging Conditional Variational AutoEncoders for Adversary Detection in the Presence of Noisy Images
- Where Classification Fails, Interpretation Rises
- FUNN: Flexible Unsupervised Neural Network
- HAWKEYE: Adversarial Example Detector for Deep Neural Networks