Using Non-invertible Data Transformations to Build Adversarial-Robust Neural Networks
arXiv:1610.01934
Abstract
Deep neural networks have proven to be quite effective in a wide variety of machine learning tasks, ranging from improved speech recognition systems to advancing the development of autonomous vehicles. However, despite their superior performance in many applications, these models have been recently shown to be susceptible to a particular type of attack possible through the generation of particular synthetic examples referred to as adversarial samples. These samples are constructed by manipulating real examples from the training data distribution in order to "fool" the original neural model, resulting in misclassification (with high confidence) of previously correctly classified samples. Addressing this weakness is of utmost importance if deep neural architectures are to be applied to critical applications, such as those in the domain of cybersecurity. In this paper, we present an analysis of this fundamental flaw lurking in all neural architectures to uncover limitations of previously proposed defense mechanisms. More importantly, we present a unifying framework for protecting deep neural models using a non-invertible data transformation--developing two adversary-resilient architectures utilizing both linear and nonlinear dimensionality reduction. Empirical results indicate that our framework provides better robustness compared to state-of-art solutions while having negligible degradation in accuracy.
References in corpus (5)
- Towards Deep Neural Network Architectures Robust to Adversarial Examples
- Learning with a Strong Adversary
- Defensive Distillation is Not Robust to Adversarial Examples
- Improving Back-Propagation by Adding an Adversarial Gradient
- Unifying Adversarial Training Algorithms with Flexible Deep Data Gradient Regularization
Cited by in corpus (5)
- Ensemble Adversarial Training: Attacks and Defenses
- Threat of Adversarial Attacks on Deep Learning in Computer Vision: A Survey
- Characterizing Audio Adversarial Examples Using Temporal Dependency
- Blocking Transferability of Adversarial Examples in Black-Box Learning Systems
- DNA Steganalysis Using Deep Recurrent Neural Networks