Adversarial Manipulation of Deep Representations
arXiv:1511.05122
Abstract
We show that the representation of an image in a deep neural network (DNN) can be manipulated to mimic those of other natural images, with only minor, imperceptible perturbations to the original image. Previous methods for generating adversarial images focused on image perturbations designed to produce erroneous class labels, while we concentrate on the internal layers of DNN representations. In this way our new class of adversarial images differs qualitatively from others. While the adversary is perceptually similar to one image, its internal representation appears remarkably similar to a different image, one from a different class, bearing little if any apparent similarity to the input; they appear generic and consistent with the space of natural images. This phenomenon raises questions about DNN representations, as well as the properties of natural images themselves.
Accepted as a conference paper at ICLR 2016
References in corpus (4)
Cited by in corpus (34)
- Robust Physical-World Attacks on Deep Learning Models
- Threat of Adversarial Attacks on Deep Learning in Computer Vision: A Survey
- advertorch v0.1: An Adversarial Robustness Toolbox based on PyTorch
- Fast Feature Fool: A data independent approach to universal adversarial perturbations
- Physical Adversarial Examples for Object Detectors
- Universal adversarial perturbations
- Optimization and Abstraction: A Synergistic Approach for Analyzing Neural Network Robustness
- Detection of Face Recognition Adversarial Attacks
- Adversarial Examples in Modern Machine Learning: A Review
- Note on Attacking Object Detectors with Adversarial Stickers
- Adversarial examples for generative models
- Learning Adversary-Resistant Deep Neural Networks
- Defense against Universal Adversarial Perturbations
- Adversarial Examples on Object Recognition: A Comprehensive Survey
- Geometric robustness of deep networks: analysis and improvement
- Identify Susceptible Locations in Medical Records via Adversarial Attacks on Deep Predictive Models
- Adversarial Attack on DL-based Massive MIMO CSI Feedback
- Built-in Vulnerabilities to Imperceptible Adversarial Perturbations
- Adversarial Attack and Defense in Deep Ranking
- The Intriguing Relation Between Counterfactual Explanations and Adversarial Examples
- SOAR: Second-Order Adversarial Regularization
- Performance Evaluation of Adversarial Attacks: Discrepancies and Solutions
- Intermediate Level Adversarial Attack for Enhanced Transferability
- Trust but Verify: An Information-Theoretic Explanation for the Adversarial Fragility of Machine Learning Systems, and a General Defense against Adversarial Attacks
- Adversarial Attacks on Co-Occurrence Features for GAN Detection
- AutoGAN: Robust Classifier Against Adversarial Attacks
- Butterfly Effect: Bidirectional Control of Classification Performance by Small Additive Perturbation
- Medical Aegis: Robust adversarial protectors for medical images
- Assessing the Adversarial Robustness of Monte Carlo and Distillation Methods for Deep Bayesian Neural Network Classification
- Training Artificial Neural Networks by Generalized Likelihood Ratio Method: Exploring Brain-like Learning to Improve Robustness
- A Useful Taxonomy for Adversarial Robustness of Neural Networks
- Can you hear me ? Sensitive comparisons of human and machine perception
- Adaptive Gradient for Adversarial Perturbations Generation
- Robust Face Verification via Disentangled Representations