Disrupting Model Training with Adversarial Shortcuts
arXiv:2106.06654
Abstract
When data is publicly released for human consumption, it is unclear how to prevent its unauthorized usage for machine learning purposes. Successful model training may be preventable with carefully designed dataset modifications, and we present a proof-of-concept approach for the image classification setting. We propose methods based on the notion of adversarial shortcuts, which encourage models to rely on non-robust signals rather than semantic features, and our experiments demonstrate that these measures successfully prevent deep learning models from achieving high accuracy on real, unmodified data examples.
References in corpus (9)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- Understanding deep learning requires rethinking generalization
- Poisoning Attacks against Support Vector Machines
- Measuring the tendency of CNNs to Learn Surface Statistical Regularities
- Transferable Clean-Label Poisoning Attacks on Deep Neural Nets
- Unlearnable Examples: Making Personal Data Unexploitable
- Adversarial Examples Make Strong Poisons
- Preventing Unauthorized Use of Proprietary Data: Poisoning for Secure Dataset Release