Poisoning and Backdooring Contrastive Learning
arXiv:2106.09667
Abstract
Multimodal contrastive learning methods like CLIP train on noisy and uncurated training datasets. This is cheaper than labeling datasets manually, and even improves out-of-distribution robustness. We show that this practice makes backdoor and poisoning attacks a significant threat. By poisoning just 0.01% of a dataset (e.g., just 300 images of the 3 million-example Conceptual Captions dataset), we can cause the model to misclassify test images by overlaying a small patch. Targeted poisoning attacks, whereby the model misclassifies a particular test input with an adversarially-desired label, are even easier requiring control of 0.0001% of the dataset (e.g., just three out of the 3 million images). Our attacks call into question whether training on noisy and uncurated Internet scrapes is desirable.
References in corpus (10)
- Learning Transferable Visual Models From Natural Language Supervision
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
- Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning
- Poisoning Attacks against Support Vector Machines
- Learning Representations by Maximizing Mutual Information Across Views
- Do ImageNet Classifiers Generalize to ImageNet?
- Security Analysis of Online Centroid Anomaly Detection
- Label-Consistent Backdoor Attacks
- A Unified Framework for Data Poisoning Attack to Graph-based Semi-supervised Learning
- Poisoning the Unlabeled Dataset of Semi-Supervised Learning
Cited by in corpus (6)
- On the Opportunities and Risks of Foundation Models
- Unsolved Problems in ML Safety
- CLOOB: Modern Hopfield Networks with InfoLOOB Outperform CLIP
- EfficientCLIP: Efficient Cross-Modal Pre-training by Ensemble Confident Learning and Language Modeling
- Better Safe Than Sorry: Preventing Delusive Adversaries with Adversarial Training
- 10 Security and Privacy Problems in Large Foundation Models