Towards Deep Neural Network Architectures Robust to Adversarial Examples
arXiv:1412.5068
Abstract
Recent work has shown deep neural networks (DNNs) to be highly susceptible to well-designed, small perturbations at the input layer, or so-called adversarial examples. Taking images as an example, such distortions are often imperceptible, but can result in 100% mis-classification for a state of the art DNN. We study the structure of adversarial examples and explore network topology, pre-processing and training strategies to improve the robustness of DNNs. We perform various experiments to assess the removability of adversarial examples by corrupting with additional noise and pre-processing with denoising autoencoders (DAEs). We find that DAEs can remove substantial amounts of the adversarial noise. How- ever, when stacking the DAE with the original DNN, the resulting network can again be attacked by new adversarial examples with even smaller distortion. As a solution, we propose Deep Contractive Network, a model with a new end-to-end training procedure that includes a smoothness penalty inspired by the contractive autoencoder (CAE). This increases the network robustness to adversarial examples, without a significant performance penalty.
References in corpus (2)
Cited by in corpus (20)
- Explaining and Harnessing Adversarial Examples
- Generating Adversarial Malware Examples for Black-Box Attacks Based on GAN
- A study of the effect of JPG compression on adversarial images
- Spectral Norm Regularization for Improving the Generalizability of Deep Learning
- NO Need to Worry about Adversarial Examples in Object Detection in Autonomous Vehicles
- Measuring the tendency of CNNs to Learn Surface Statistical Regularities
- Simple Black-Box Adversarial Perturbations for Deep Networks
- Robustness of classifiers: from adversarial to random noise
- Adversarial Examples that Fool Detectors
- Biologically inspired protection of deep networks from adversarial attacks
- Blocking Transferability of Adversarial Examples in Black-Box Learning Systems
- Adversarial Example Defenses: Ensembles of Weak Defenses are not Strong
- APE-GAN: Adversarial Perturbation Elimination with GAN
- Exploring the Space of Black-box Attacks on Deep Neural Networks
- Adversarial Images for Variational Autoencoders
- Standard detectors aren't (currently) fooled by physical adversarial stop signs
- Using Non-invertible Data Transformations to Build Adversarial-Robust Neural Networks
- Neural Networks Regularization Through Class-wise Invariant Representation Learning
- Butterfly Effect: Bidirectional Control of Classification Performance by Small Additive Perturbation
- Adequacy of the Gradient-Descent Method for Classifier Evasion Attacks