Early Methods for Detecting Adversarial Images
arXiv:1608.00530
Abstract
Many machine learning classifiers are vulnerable to adversarial perturbations. An adversarial perturbation modifies an input to change a classifier's prediction without causing the input to seem substantially different to human perception. We deploy three methods to detect adversarial images. Adversaries trying to bypass our detectors must make the adversarial image less pathological or they will fail trying. Our best detection method reveals that adversarial images place abnormal emphasis on the lower-ranked principal components from PCA. Other detectors and a colorful saliency map are in an appendix.
ICLR 2017 Workshop Contribution
Cited by in corpus (11)
- Two Souls in an Adversarial Image: Towards Universal Adversarial Example Detection using Multi-view Inconsistency
- Connecting the Dots: Detecting Adversarial Perturbations Using Context Inconsistency
- Cassandra: Detecting Trojaned Networks from Adversarial Perturbations
- ExAD: An Ensemble Approach for Explanation-based Adversarial Detection
- Long-term Cross Adversarial Training: A Robust Meta-learning Method for Few-shot Classification Tasks
- Moving Target Defense for Deep Visual Sensing against Adversarial Examples
- Efficient detection of adversarial images
- Effect of backdoor attacks over the complexity of the latent space distribution
- Using Anomaly Feature Vectors for Detecting, Classifying and Warning of Outlier Adversarial Examples
- Detecting Adversaries, yet Faltering to Noise? Leveraging Conditional Variational AutoEncoders for Adversary Detection in the Presence of Noisy Images
- Attribution of Gradient Based Adversarial Attacks for Reverse Engineering of Deceptions