Generative Adversarial Trainer: Defense to Adversarial Perturbations with GAN
arXiv:1705.03387
Abstract
We propose a novel technique to make neural network robust to adversarial examples using a generative adversarial network. We alternately train both classifier and generator networks. The generator network generates an adversarial perturbation that can easily fool the classifier network by using a gradient of each image. Simultaneously, the classifier network is trained to classify correctly both original and adversarial images generated by the generator. These procedures help the classifier network to become more robust to adversarial perturbations. Furthermore, our adversarial training framework efficiently reduces overfitting and outperforms other regularization methods such as Dropout. We applied our method to supervised learning for CIFAR datasets, and experimantal results show that our method significantly lowers the generalization error of the network. To the best of our knowledge, this is the first method which uses GAN to improve supervised learning.
References in corpus (2)
Cited by in corpus (34)
- IDSGAN: Generative Adversarial Networks for Attack Generation against Intrusion Detection
- A Review on Generative Adversarial Networks: Algorithms, Theory, and Applications
- Adversarial Examples: Opportunities and Challenges
- Threat of Adversarial Attacks on Deep Learning in Computer Vision: A Survey
- Motivating the Rules of the Game for Adversarial Example Research
- Constructing Unrestricted Adversarial Examples with Generative Models
- The Robust Manifold Defense: Adversarial Training using Generative Models
- Adversarial Training against Location-Optimized Adversarial Patches
- Adversarial Examples - A Complete Characterisation of the Phenomenon
- High Frequency Component Helps Explain the Generalization of Convolutional Neural Networks
- Defense Methods Against Adversarial Examples for Recurrent Neural Networks
- Generative Adversarial Networks: A Survey Towards Private and Secure Applications
- Improving adversarial robustness of deep neural networks by using semantic information
- Medical Image Harmonization Using Deep Learning Based Canonical Mapping: Toward Robust and Generalizable Learning in Imaging
- FenceBox: A Platform for Defeating Adversarial Examples with Data Augmentation Techniques
- Mitigating Advanced Adversarial Attacks with More Advanced Gradient Obfuscation Techniques
- Adversarial Examples on Object Recognition: A Comprehensive Survey
- Efficient Adversarial Training with Transferable Adversarial Examples
- Retrieval-Augmented Convolutional Neural Networks for Improved Robustness against Adversarial Examples
- Improving Global Adversarial Robustness Generalization With Adversarially Trained GAN
- ATMPA: Attacking Machine Learning-based Malware Visualization Detection Methods via Adversarial Examples
- Distortion Agnostic Deep Watermarking
- Model-Based Robust Deep Learning: Generalizing to Natural, Out-of-Distribution Data
- Improved Adversarial Training via Learned Optimizer
- Distributional Robustness with IPMs and links to Regularization and GANs
- Adversarial Attack Type I: Cheat Classifiers by Significant Changes
- FineFool: Fine Object Contour Attack via Attention
- Implicit Generative Modeling of Random Noise during Training for Adversarial Robustness
- Using Randomness to Improve Robustness of Machine-Learning Models Against Evasion Attacks
- Generative Models for Security: Attacks, Defenses, and Opportunities
- Towards Natural Robustness Against Adversarial Examples
- Orthogonal Deep Models As Defense Against Black-Box Attacks
- FUNN: Flexible Unsupervised Neural Network
- Evolution Attack On Neural Networks