Adversarial Transformation Networks: Learning to Generate Adversarial Examples
arXiv:1703.09387
Abstract
Multiple different approaches of generating adversarial examples have been proposed to attack deep neural networks. These approaches involve either directly computing gradients with respect to the image pixels, or directly solving an optimization on the image pixels. In this work, we present a fundamentally new method for generating adversarial examples that is fast to execute and provides exceptional diversity of output. We efficiently train feed-forward neural networks in a self-supervised manner to generate adversarial examples against a target network or set of networks. We call such a network an Adversarial Transformation Network (ATN). ATNs are trained to generate adversarial examples that minimally modify the classifier's outputs given the original input, while constraining the new classification to match an adversarial target class. We present methods to train ATNs and analyze their effectiveness targeting a variety of MNIST classifiers as well as the latest state-of-the-art ImageNet classifier Inception ResNet v2.
Cited by in corpus (28)
- UPSET and ANGRI : Breaking High Performance Image Classifiers
- Model Patching: Closing the Subgroup Performance Gap with Data Augmentation
- Triple Wins: Boosting Accuracy, Robustness and Efficiency Together by Enabling Input-Adaptive Inference
- Generative Adversarial Networks: A Survey Towards Private and Secure Applications
- Stabilized Medical Image Attacks
- Security and Privacy for Artificial Intelligence: Opportunities and Challenges
- Adversarial Examples on Object Recognition: A Comprehensive Survey
- Geometric robustness of deep networks: analysis and improvement
- A Closer Look at Data Bias in Neural Extractive Summarization Models
- GAP++: Learning to generate target-conditioned adversarial examples
- Learning from Higher-Layer Feature Visualizations
- DAmageNet: A Universal Adversarial Dataset
- Transferable Sparse Adversarial Attack
- A geometry-inspired decision-based attack
- Distortion Agnostic Deep Watermarking
- -ML: Mitigating Adversarial Examples via Ensembles of Topologically Manipulated Classifiers
- Multi-Task Adversarial Attack
- Butterfly Effect: Bidirectional Control of Classification Performance by Small Additive Perturbation
- Improving the Transferability of Adversarial Examples with New Iteration Framework and Input Dropout
- Prototype-supervised Adversarial Network for Targeted Attack of Deep Hashing
- Moving Target Defense for Deep Visual Sensing against Adversarial Examples
- Thinking Outside the Pool: Active Training Image Creation for Relative Attributes
- Techniques for Adversarial Examples Threatening the Safety of Artificial Intelligence Based Systems
- Adversarial Black-Box Attacks On Text Classifiers Using Multi-Objective Genetic Optimization Guided By Deep Networks
- Model-Agnostic Meta-Attack: Towards Reliable Evaluation of Adversarial Robustness
- Fast Local Attack: Generating Local Adversarial Examples for Object Detectors
- Defending Against Adversarial Attacks Using Random Forests
- Evolution Attack On Neural Networks