Poisoning the Unlabeled Dataset of Semi-Supervised Learning
arXiv:2105.01622
Abstract
Semi-supervised machine learning models learn from a (small) set of labeled training examples, and a (large) set of unlabeled training examples. State-of-the-art models can reach within a few percentage points of fully-supervised training, while requiring 100x less labeled data. We study a new class of vulnerabilities: poisoning attacks that modify the unlabeled dataset. In order to be useful, unlabeled datasets are given strictly less review than labeled datasets, and adversaries can therefore poison them easily. By inserting maliciously-crafted unlabeled examples totaling just 0.1% of the dataset size, we can manipulate a model trained on this poisoned dataset to misclassify arbitrary examples at test time (as any desired label). Our attacks are highly effective across datasets and semi-supervised learning methods. We find that more accurate methods (thus more likely to be used) are significantly more vulnerable to poisoning attacks, and as such better training methods are unlikely to prevent this attack. To counter this we explore the space of defenses, and propose two methods that mitigate our attack.
References in corpus (8)
- Improved Regularization of Convolutional Neural Networks with Cutout
- Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning
- Poisoning Attacks against Support Vector Machines
- Do ImageNet Classifiers Generalize to ImageNet?
- Security Analysis of Online Centroid Anomaly Detection
- On the Effectiveness of Mitigating Data Poisoning Attacks with Gradient Shaping
- Label-Consistent Backdoor Attacks
- A Unified Framework for Data Poisoning Attack to Graph-based Semi-supervised Learning
Cited by in corpus (6)
- On the Opportunities and Risks of Foundation Models
- Multi-SpacePhish: Extending the Evasion-space of Adversarial Attacks against Phishing Website Detectors using Machine Learning
- Wild Networks: Exposure of 5G Network Infrastructures to Adversarial Examples
- Poisoning and Backdooring Contrastive Learning
- Poisoning Attacks to Local Differential Privacy Protocols for Key-Value Data
- Better Safe Than Sorry: Preventing Delusive Adversaries with Adversarial Training