On the (Statistical) Detection of Adversarial Examples
arXiv:1702.06280
Abstract
Machine Learning (ML) models are applied in a variety of tasks such as network intrusion detection or Malware classification. Yet, these models are vulnerable to a class of malicious inputs known as adversarial examples. These are slightly perturbed inputs that are classified incorrectly by the ML model. The mitigation of these adversarial inputs remains an open problem. As a step towards understanding adversarial examples, we show that they are not drawn from the same distribution than the original data, and can thus be detected using statistical tests. Using thus knowledge, we introduce a complimentary approach to identify specific inputs that are adversarial. Specifically, we augment our ML model with an additional output, in which the model is trained to classify all adversarial inputs. We evaluate our approach on multiple adversarial example crafting methods (including the fast gradient sign and saliency map methods) with several datasets. The statistical test flags sample sets containing adversarial inputs confidently at sample sizes between 10 and 100 data points. Furthermore, our augmented model either detects adversarial examples as outliers with high accuracy (> 80%) or increases the adversary's cost - the perturbation added - by more than 150%. In this way, we show that statistical properties of adversarial examples are essential to their detection.
13 pages, 4 figures, 5 tables. New version: improved writing, incorporating external feedback
References in corpus (3)
Cited by in corpus (169)
- Feature Squeezing: Detecting Adversarial Examples in Deep Neural Networks
- ZOO: Zeroth Order Optimization based Black-box Attacks to Deep Neural Networks without Training Substitute Models
- Characterizing Adversarial Subspaces Using Local Intrinsic Dimensionality
- Adversarial Frontier Stitching for Remote Neural Network Watermarking
- PixelDefend: Leveraging Generative Models to Understand and Defend against Adversarial Examples
- Adversarial Examples: Opportunities and Challenges
- Adversarial Examples: Attacks and Defenses for Deep Learning
- MagNet: a Two-Pronged Defense against Adversarial Examples
- DeepTest: Automated Testing of Deep-Neural-Network-driven Autonomous Cars
- Adversarial Weight Perturbation Helps Robust Generalization
- A General Framework for Adversarial Examples with Objectives
- Adversarial Example Detection for DNN Models: A Review and Experimental Comparison
- Local Model Poisoning Attacks to Byzantine-Robust Federated Learning
- Threat of Adversarial Attacks on Deep Learning in Computer Vision: A Survey
- Motivating the Rules of the Game for Adversarial Example Research
- The Robust Manifold Defense: Adversarial Training using Generative Models
- A Review of Adversarial Attack and Defense for Classification Methods
- Extending Defensive Distillation
- Adversarial Example Defenses: Ensembles of Weak Defenses are not Strong
- Malware Makeover: Breaking ML-based Static Analysis by Modifying Executable Bytes
- A Simple Explanation for the Existence of Adversarial Examples with Small Hamming Distance
- Reverse KL-Divergence Training of Prior Networks: Improved Uncertainty and Adversarial Robustness
- EAD: Elastic-Net Attacks to Deep Neural Networks via Adversarial Examples
- Security and Privacy Issues in Deep Learning
- Detection of Face Recognition Adversarial Attacks
- Detecting and Diagnosing Adversarial Images with Class-Conditional Capsule Reconstructions
- A Survey of Game Theoretic Approaches for Adversarial Machine Learning in Cybersecurity Tasks
- WAF-A-MoLE: Evading Web Application Firewalls through Adversarial Machine Learning
- Adversarial Examples in Modern Machine Learning: A Review
- Towards Robust Neural Networks via Random Self-ensemble
- Adversarial Examples - A Complete Characterisation of the Phenomenon
- Detecting Adversarial Samples for Deep Neural Networks through Mutation Testing
- A Survey on Open Set Recognition
- Why is the Mahalanobis Distance Effective for Anomaly Detection?
- Arms Race in Adversarial Malware Detection: A Survey
- Malytics: A Malware Detection Scheme
- Defense Methods Against Adversarial Examples for Recurrent Neural Networks
- Curse of Dimensionality on Randomized Smoothing for Certifiable Robustness
- Confidence-Calibrated Adversarial Training: Generalizing to Unseen Attacks
- Detecting Adversarial Attacks on Neural Network Policies with Visual Foresight
- Daedalus: Breaking Non-Maximum Suppression in Object Detection via Adversarial Examples
- PRADA: Protecting against DNN Model Stealing Attacks
- Security for Machine Learning-based Systems: Attacks and Challenges during Training and Inference
- Adversarial Robustness via Fisher-Rao Regularization
- Are Generative Classifiers More Robust to Adversarial Attacks?
- Understanding the One-Pixel Attack: Propagation Maps and Locality Analysis
- Adversarial Learning in Statistical Classification: A Comprehensive Review of Defenses Against Attacks
- Minimax Robust Detection: Classic Results and Recent Advances
- When Explainability Meets Adversarial Learning: Detecting Adversarial Examples using SHAP Signatures
- White Paper Machine Learning in Certified Systems
- Attack and Defense of Dynamic Analysis-Based, Adversarial Neural Malware Classification Models
- PAC-learning in the presence of evasion adversaries
- Attacking Optical Character Recognition (OCR) Systems with Adversarial Watermarks
- RANDOM MASK: Towards Robust Convolutional Neural Networks
- ML-LOO: Detecting Adversarial Examples with Feature Attribution
- A Provable Defense for Deep Residual Networks
- Deflecting Adversarial Attacks
- Natural and Adversarial Error Detection using Invariance to Image Transformations
- Adversarial Attacks for Tabular Data: Application to Fraud Detection and Imbalanced Data
- Breaking Transferability of Adversarial Samples with Randomness
- RAID: Randomized Adversarial-Input Detection for Neural Networks
- Security and Privacy for Artificial Intelligence: Opportunities and Challenges
- Two Souls in an Adversarial Image: Towards Universal Adversarial Example Detection using Multi-view Inconsistency
- DaST: Data-free Substitute Training for Adversarial Attacks
- Adversarial Example Games
- Towards Robust Detection of Adversarial Examples
- Improving Network Robustness against Adversarial Attacks with Compact Convolution
- Noise Sensitivity-Based Energy Efficient and Robust Adversary Detection in Neural Networks
- Generic Black-Box End-to-End Attack Against State of the Art API Call Based Malware Classifiers
- ATHENA: A Framework based on Diverse Weak Defenses for Building Adversarial Defense
- Defence against adversarial attacks using classical and quantum-enhanced Boltzmann machines
- SoK: Machine Learning Governance
- Adversarial Examples on Object Recognition: A Comprehensive Survey
- Detecting Adversarial Examples Is (Nearly) As Hard As Classifying Them
- Perturbation Analysis of Gradient-based Adversarial Attacks
- The Taboo Trap: Behavioural Detection of Adversarial Samples
- Generating Semantic Adversarial Examples via Feature Manipulation
- Certifying Confidence via Randomized Smoothing
- Fooling Vision and Language Models Despite Localization and Attention Mechanism
- Adversarial Robustness Assessment: Why both and Attacks Are Necessary
- Anomalous Example Detection in Deep Learning: A Survey
- Input Validation for Neural Networks via Runtime Local Robustness Verification
- Using Undervolting as an On-Device Defense Against Adversarial Machine Learning Attacks
- Controlling Over-generalization and its Effect on Adversarial Examples Generation and Detection
- Effective and Robust Detection of Adversarial Examples via Benford-Fourier Coefficients
- Detecting Adversarial Examples via Key-based Network
- Detection based Defense against Adversarial Examples from the Steganalysis Point of View
- Better the Devil you Know: An Analysis of Evasion Attacks using Out-of-Distribution Adversarial Examples
- Runtime Monitoring Neuron Activation Patterns
- LiBRe: A Practical Bayesian Approach to Adversarial Detection
- Adversarial Attacks and Defenses: An Interpretation Perspective
- Mitigating Evasion Attacks to Deep Neural Networks via Region-based Classification
- Using Depth for Pixel-Wise Detection of Adversarial Attacks in Crowd Counting
- Attacking and Defending Machine Learning Applications of Public Cloud
- Built-in Vulnerabilities to Imperceptible Adversarial Perturbations
- Detecting Adversarial Examples from Sensitivity Inconsistency of Spatial-Transform Domain
- Defending against Adversarial Attack towards Deep Neural Networks via Collaborative Multi-task Training
- Subset Scanning Over Neural Network Activations
- Stateful Detection of Black-Box Adversarial Attacks
- -ML: Mitigating Adversarial Examples via Ensembles of Topologically Manipulated Classifiers
- Cassandra: Detecting Trojaned Networks from Adversarial Perturbations
- TREATED:Towards Universal Defense against Textual Adversarial Attacks
- Adversarial Classification: Necessary conditions and geometric flows
- Random Directional Attack for Fooling Deep Neural Networks
- DNA Steganalysis Using Deep Recurrent Neural Networks
- Universal Adversarial Perturbations Through the Lens of Deep Steganography: Towards A Fourier Perspective
- Input Hessian Regularization of Neural Networks
- Automated Discovery of Adaptive Attacks on Adversarial Defenses
- Prior Networks for Detection of Adversarial Attacks
- Adversarial Robustness in Deep Learning: Attacks on Fragile Neurons
- Picket: Guarding Against Corrupted Data in Tabular Data during Learning and Inference
- Analytical Moment Regularizer for Gaussian Robust Networks
- Detecting Adversarial Perturbations with Saliency
- Omni: Automated Ensemble with Unexpected Models against Adversarial Evasion Attack
- TESDA: Transform Enabled Statistical Detection of Attacks in Deep Neural Networks
- Technologies for Trustworthy Machine Learning: A Survey in a Socio-Technical Context
- Policy Smoothing for Provably Robust Reinforcement Learning
- Improving Model Robustness with Latent Distribution Locally and Globally
- Most ReLU Networks Suffer from Adversarial Perturbations
- ExAD: An Ensemble Approach for Explanation-based Adversarial Detection
- A New Angle on L2 Regularization
- DAFAR: Defending against Adversaries by Feedback-Autoencoder Reconstruction
- Identifying Untrustworthy Predictions in Neural Networks by Geometric Gradient Analysis
- Detecting Adversarial Patches with Class Conditional Reconstruction Networks
- MAD-VAE: Manifold Awareness Defense Variational Autoencoder
- AdvMS: A Multi-source Multi-cost Defense Against Adversarial Attacks
- Detecting and Correcting Adversarial Images Using Image Processing Operations
- Adversarial Embedding: A robust and elusive Steganography and Watermarking technique
- A Computationally Efficient Method for Defending Adversarial Deep Learning Attacks
- Long-term Cross Adversarial Training: A Robust Meta-learning Method for Few-shot Classification Tasks
- Moving Target Defense for Deep Visual Sensing against Adversarial Examples
- Unsupervised Detection of Adversarial Examples with Model Explanations
- An Empirical Study on the Relation between Network Interpretability and Adversarial Robustness
- Using Anomaly Feature Vectors for Detecting, Classifying and Warning of Outlier Adversarial Examples
- When the Guard failed the Droid: A case study of Android malware
- A Survey on Assessing the Generalization Envelope of Deep Neural Networks: Predictive Uncertainty, Out-of-distribution and Adversarial Samples
- Perceptual Deep Neural Networks: Adversarial Robustness through Input Recreation
- MixDefense: A Defense-in-Depth Framework for Adversarial Example Detection Based on Statistical and Semantic Analysis
- DetectX -- Adversarial Input Detection using Current Signatures in Memristive XBar Arrays
- Immuno-mimetic Deep Neural Networks (Immuno-Net)
- Defending Adversarial Attacks via Semantic Feature Manipulation
- Mitigation of Adversarial Attacks through Embedded Feature Selection
- Models of Computational Profiles to Study the Likelihood of DNN Metamorphic Test Cases
- Identification of Attack-Specific Signatures in Adversarial Examples
- Clipping free attacks against artificial neural networks
- Out-distribution training confers robustness to deep neural networks
- On Lyapunov exponents and adversarial perturbation
- Towards Imperceptible Query-limited Adversarial Attacks with Perceptual Feature Fidelity Loss
- Detecting Adversarial Samples Using Density Ratio Estimates
- Exploiting the Sensitivity of Adversarial Examples to Erase-and-Restore
- HAWKEYE: Adversarial Example Detector for Deep Neural Networks
- Deep Bayesian Image Set Classification: A Defence Approach against Adversarial Attacks
- An Explainable Adversarial Robustness Metric for Deep Learning Neural Networks
- Deep Adversarially-Enhanced k-Nearest Neighbors
- Attribution of Gradient Based Adversarial Attacks for Reverse Engineering of Deceptions
- Adversarial Attacks with Time-Scale Representations
- Decoder-free Robustness Disentanglement without (Additional) Supervision
- A Primer on Multi-Neuron Relaxation-based Adversarial Robustness Certification
- A fast and effective kernel two-sample test for large-scale data
- Oriole: Thwarting Privacy against Trustworthy Deep Learning Models
- Denoising and Verification Cross-Layer Ensemble Against Black-box Adversarial Attacks
- Interpretable BoW Networks for Adversarial Example Detection
- Adaptive Gradient for Adversarial Perturbations Generation
- Where Classification Fails, Interpretation Rises
- A novel network training approach for open set image recognition
- Learning to Separate Clusters of Adversarial Representations for Robust Adversarial Detection
- Learning to Disentangle Robust and Vulnerable Features for Adversarial Detection
- Adversarial Attack Attribution: Discovering Attributable Signals in Adversarial ML Attacks
- Adversarial Imitation Attack