Towards Robust Neural Networks via Random Self-ensemble
arXiv:1712.00673
Abstract
Recent studies have revealed the vulnerability of deep neural networks: A small adversarial perturbation that is imperceptible to human can easily make a well-trained deep neural network misclassify. This makes it unsafe to apply neural networks in security-critical applications. In this paper, we propose a new defense algorithm called Random Self-Ensemble (RSE) by combining two important concepts: {\bf randomness} and {\bf ensemble}. To protect a targeted model, RSE adds random noise layers to the neural network to prevent the strong gradient-based attacks, and ensembles the prediction over random noises to stabilize the performance. We show that our algorithm is equivalent to ensemble an infinite number of noisy models without any additional memory overhead, and the proposed training procedure based on noisy stochastic gradient descent can ensure the ensemble model has a good predictive capability. Our algorithm significantly outperforms previous defense techniques on real data sets. For instance, on CIFAR-10 with VGG network (which has 92\% accuracy without any attack), under the strong C\&W attack within a certain distortion tolerance, the accuracy of unprotected model drops to less than 10\%, the best previous defense technique has accuracy, while our method still has prediction accuracy under the same level of attack. Finally, our method is simple and easy to integrate into any neural network.
ECCV 2018 camera ready
References in corpus (8)
- Obfuscated Gradients Give a False Sense of Security: Circumventing Defenses to Adversarial Examples
- Characterizing Adversarial Subspaces Using Local Intrinsic Dimensionality
- Countering Adversarial Images using Input Transformations
- Adversarial Machine Learning at Scale
- PixelDefend: Leveraging Generative Models to Understand and Defend against Adversarial Examples
- Stochastic Activation Pruning for Robust Adversarial Defense
- Evaluating the Robustness of Neural Networks: An Extreme Value Theory Approach
- Ensemble Methods as a Defense to Adversarial Perturbations Against Deep Neural Networks
Cited by in corpus (14)
- Adversarial Machine Learning in Image Classification: A Survey Towards the Defender's Perspective
- Evaluating the Robustness of Neural Networks: An Extreme Value Theory Approach
- Towards Fast Computation of Certified Robustness for ReLU Networks
- Adv-BNN: Improved Adversarial Defense through Robust Bayesian Neural Network
- Improving Adversarial Robustness of Ensembles with Diversity Training
- Adversarial Machine Learning And Speech Emotion Recognition: Utilizing Generative Adversarial Networks For Robustness
- Robust Sparse Regularization: Simultaneously Optimizing Neural Network Robustness and Compactness
- Second-Order Provable Defenses against Adversarial Attacks
- Optimal Transport Classifier: Defending Against Adversarial Attacks by Regularized Deep Embedding
- Robustness Certificates Against Adversarial Examples for ReLU Networks
- Efficient Bidirectional Neural Machine Translation
- AdvMS: A Multi-source Multi-cost Defense Against Adversarial Attacks
- An Empirical Investigation of Randomized Defenses against Adversarial Attacks
- Closing the Gap: Achieving Better Accuracy-Robustness Tradeoffs against Query-Based Attacks