Better Safe Than Sorry: Preventing Delusive Adversaries with Adversarial Training
arXiv:2102.04716
Abstract
Delusive attacks aim to substantially deteriorate the test accuracy of the learning model by slightly perturbing the features of correctly labeled training examples. By formalizing this malicious attack as finding the worst-case training data within a specific -Wasserstein ball, we show that minimizing adversarial risk on the perturbed data is equivalent to optimizing an upper bound of natural risk on the original data. This implies that adversarial training can serve as a principled defense against delusive attacks. Thus, the test accuracy decreased by delusive attacks can be largely recovered by adversarial training. To further understand the internal mechanism of the defense, we disclose that adversarial training can resist the delusive perturbations by preventing the learner from overly relying on non-robust features in a natural setting. Finally, we complement our theoretical findings with a set of experiments on popular benchmark datasets, which show that the defense withstands six different practical attacks. Both theoretical and empirical results vote for adversarial training when confronted with delusive adversaries.
NeurIPS 2021
References in corpus (28)
- Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning
- Theoretically Principled Trade-off between Robustness and Accuracy
- Poisoning Attacks against Support Vector Machines
- On the Convergence and Robustness of Adversarial Training
- Self-supervised Visual Feature Learning with Deep Neural Networks: A Survey
- Neural Attention Distillation: Erasing Backdoor Triggers from Deep Neural Networks
- Transferable Clean-Label Poisoning Attacks on Deep Neural Nets
- Interpreting Adversarially Trained Convolutional Neural Networks
- On the Effectiveness of Mitigating Data Poisoning Attacks with Gradient Shaping
- Label-Consistent Backdoor Attacks
- Witches' Brew: Industrial Scale Data Poisoning via Gradient Matching
- Anti-Backdoor Learning: Training Clean Models on Poisoned Data
- A Unified Framework for Data Poisoning Attack to Graph-based Semi-supervised Learning
- To be Robust or to be Fair: Towards Fairness in Adversarial Training
- LowKey: Leveraging Adversarial Attacks to Protect Social Media Users from Facial Recognition
- Poisoning the Unlabeled Dataset of Semi-Supervised Learning
- Poisoning and Backdooring Contrastive Learning
- On the effectiveness of adversarial training against common corruptions
- Maximum Mean Discrepancy Test is Aware of Adversarial Attacks
- Understanding the Interaction of Adversarial Training with Noisy Labels
- Exploring Memorization in Adversarial Training
- Preventing Unauthorized Use of Proprietary Data: Poisoning for Secure Dataset Release
- What Do Deep Nets Learn? Class-wise Patterns Revealed in the Input Space
- Regularization Helps with Mitigating Poisoning Attacks: Distributionally-Robust Machine Learning Using the Wasserstein Distance
- Consistency Regularization for Adversarial Robustness
- Disrupting Model Training with Adversarial Shortcuts
- Black-box Detection of Backdoor Attacks with Limited Information and Data
- On Evaluating Neural Network Backdoor Defenses