Just How Toxic is Data Poisoning? A Unified Benchmark for Backdoor and Data Poisoning Attacks
arXiv:2006.12557
Abstract
Data poisoning and backdoor attacks manipulate training data in order to cause models to fail during inference. A recent survey of industry practitioners found that data poisoning is the number one concern among threats ranging from model stealing to adversarial attacks. However, it remains unclear exactly how dangerous poisoning methods are and which ones are more effective considering that these methods, even ones with identical objectives, have not been tested in consistent or realistic settings. We observe that data poisoning and backdoor attacks are highly sensitive to variations in the testing setup. Moreover, we find that existing methods may not generalize to realistic settings. While these existing works serve as valuable prototypes for data poisoning, we apply rigorous tests to determine the extent to which we should fear them. In order to promote fair comparison in future work, we develop standardized benchmarks for data poisoning and backdoor attacks.
19 pages, 4 figures
References in corpus (9)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning
- Poisoning Attacks against Support Vector Machines
- Support Vector Machines under Adversarial Label Contamination
- Data Poisoning Attacks on Factorization-Based Collaborative Filtering
- Transferable Clean-Label Poisoning Attacks on Deep Neural Nets
- MetaPoison: Practical General-purpose Clean-label Data Poisoning
- Witches' Brew: Industrial Scale Data Poisoning via Gradient Matching
- Don't Trigger Me! A Triggerless Backdoor Attack Against Deep Neural Networks
Cited by in corpus (14)
- Backdoor Attacks and Countermeasures on Deep Learning: A Comprehensive Review
- Rethinking the Trigger of Backdoor Attack
- Witches' Brew: Industrial Scale Data Poisoning via Gradient Matching
- DP-InstaHide: Provably Defusing Poisoning and Backdoor Attacks with Differentially Private Data Augmentations
- Defending Against Backdoor Attacks in Natural Language Generation
- Data Poisoning Vulnerabilities Across Healthcare AI Architectures: A Security Threat Analysis
- Sleeper Agent: Scalable Hidden Trigger Backdoors for Neural Networks Trained from Scratch
- Better Safe Than Sorry: Preventing Delusive Adversaries with Adversarial Training
- Advancing the Research and Development of Assured Artificial Intelligence and Machine Learning Capabilities
- Towards Audit Requirements for AI-based Systems in Mobility Applications
- Disrupting Model Training with Adversarial Shortcuts
- When AI reviews science: Can we trust the referee?
- Broadly Applicable Targeted Data Sample Omission Attacks
- Correlation between image quality metrics of magnetic resonance images and the neural network segmentation accuracy