Jacobian Adversarially Regularized Networks for Robustness
arXiv:1912.10185
Abstract
Adversarial examples are crafted with imperceptible perturbations with the intent to fool neural networks. Against such attacks, adversarial training and its variants stand as the strongest defense to date. Previous studies have pointed out that robust models that have undergone adversarial training tend to produce more salient and interpretable Jacobian matrices than their non-robust counterparts. A natural question is whether a model trained with an objective to produce salient Jacobian can result in better robustness. This paper answers this question with affirmative empirical results. We propose Jacobian Adversarially Regularized Networks (JARN) as a method to optimize the saliency of a classifier's Jacobian by adversarially regularizing the model's Jacobian to resemble natural training images. Image classifiers trained with JARN show improved robust accuracy compared to standard models on the MNIST, SVHN and CIFAR-10 datasets, uncovering a new angle to boost robustness without using adversarial training examples.
ICLR 2020 Camera Ready
References in corpus (5)
- Theoretically Principled Trade-off between Robustness and Accuracy
- On Evaluating Adversarial Robustness
- Improving the Adversarial Robustness and Interpretability of Deep Neural Networks by Regularizing their Input Gradients
- Adversarial Examples Are a Natural Consequence of Test Error in Noise
- PROVEN: Certifying Robustness of Neural Networks with a Probabilistic Approach