Multi-way Encoding for Robustness
arXiv:1906.02033
Abstract
Deep models are state-of-the-art for many computer vision tasks including image classification and object detection. However, it has been shown that deep models are vulnerable to adversarial examples. We highlight how one-hot encoding directly contributes to this vulnerability and propose breaking away from this widely-used, but highly-vulnerable mapping. We demonstrate that by leveraging a different output encoding, multi-way encoding, we decorrelate source and target models, making target models more secure. Our approach makes it more difficult for adversaries to find useful gradients for generating adversarial attacks. We present robustness for black-box and white-box attacks on four benchmark datasets: MNIST, CIFAR-10, CIFAR-100, and SVHN. The strength of our approach is also presented in the form of an attack for model watermarking, raising challenges in detecting stolen models.
Accepted at WACV 2020
References in corpus (11)
- Ensemble Adversarial Training: Attacks and Defenses
- Obfuscated Gradients Give a False Sense of Security: Circumventing Defenses to Adversarial Examples
- Delving into Transferable Adversarial Examples and Black-box Attacks
- On Evaluating Adversarial Robustness
- Beyond One-hot Encoding: lower dimensional target embedding
- Adversarial Risk and the Dangers of Evaluating Against Weak Attacks
- Defensive Distillation is Not Robust to Adversarial Examples
- Towards the first adversarially robust neural network model on MNIST
- Adversarial Examples Are a Natural Consequence of Test Error in Noise
- Evaluating and Understanding the Robustness of Adversarial Logit Pairing
- NATTACK: Learning the Distributions of Adversarial Examples for an Improved Black-Box Attack on Deep Neural Networks