Optimal Transport Classifier: Defending Against Adversarial Attacks by Regularized Deep Embedding
arXiv:1811.07950
Abstract
Recent studies have demonstrated the vulnerability of deep convolutional neural networks against adversarial examples. Inspired by the observation that the intrinsic dimension of image data is much smaller than its pixel space dimension and the vulnerability of neural networks grows with the input dimension, we propose to embed high-dimensional input images into a low-dimensional space to perform classification. However, arbitrarily projecting the input images to a low-dimensional space without regularization will not improve the robustness of deep neural networks. Leveraging optimal transport theory, we propose a new framework, Optimal Transport Classifier (OT-Classifier), and derive an objective that minimizes the discrepancy between the distribution of the true label and the distribution of the OT-Classifier output. Experimental results on several benchmark datasets show that, our proposed framework achieves state-of-the-art performance against strong adversarial attack methods.
9 pages
References in corpus (7)
- Explaining and Harnessing Adversarial Examples
- ZOO: Zeroth Order Optimization based Black-box Attacks to Deep Neural Networks without Training Substitute Models
- Obfuscated Gradients Give a False Sense of Security: Circumventing Defenses to Adversarial Examples
- Defense-GAN: Protecting Classifiers Against Adversarial Attacks Using Generative Models
- Mitigating Adversarial Effects Through Randomization
- Wasserstein Auto-Encoders
- Towards Robust Neural Networks via Random Self-ensemble