PatchCleanser: Certifiably Robust Defense against Adversarial Patches for Any Image Classifier
arXiv:2108.09135
Abstract
The adversarial patch attack against image classification models aims to inject adversarially crafted pixels within a restricted image region (i.e., a patch) for inducing model misclassification. This attack can be realized in the physical world by printing and attaching the patch to the victim object; thus, it imposes a real-world threat to computer vision systems. To counter this threat, we design PatchCleanser as a certifiably robust defense against adversarial patches. In PatchCleanser, we perform two rounds of pixel masking on the input image to neutralize the effect of the adversarial patch. This image-space operation makes PatchCleanser compatible with any state-of-the-art image classifier for achieving high accuracy. Furthermore, we can prove that PatchCleanser will always predict the correct class labels on certain images against any adaptive white-box attacker within our threat model, achieving certified robustness. We extensively evaluate PatchCleanser on the ImageNet, ImageNette, CIFAR-10, CIFAR-100, SVHN, and Flowers-102 datasets and demonstrate that our defense achieves similar clean accuracy as state-of-the-art classification models and also significantly improves certified robustness from prior works. Remarkably, PatchCleanser achieves 83.9% top-1 clean accuracy and 62.1% top-1 certified robust accuracy against a 2%-pixel square patch anywhere on the image for the 1000-class ImageNet dataset.
USENIX Security Symposium 2022; extended technical report
References in corpus (10)
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- Improved Regularization of Convolutional Neural Networks with Cutout
- MLP-Mixer: An all-MLP Architecture for Vision
- Certified Adversarial Robustness via Randomized Smoothing
- On Detecting Adversarial Perturbations
- Adversarial YOLO: Defense Human Detection Patch Attacks via Detecting Adversarial Patches
- PatchGuard++: Efficient Provable Attack Detection against Adversarial Patches
- Efficient Certified Defenses Against Patch Attacks on Image Classifiers
- A Real-time Defense against Website Fingerprinting Attacks
- Detecting Adversarial Examples Is (Nearly) As Hard As Classifying Them