I Know What You Trained Last Summer: A Survey on Stealing Machine Learning Models and Defences
arXiv:2206.08451 · doi:10.1145/3595292
Abstract
Machine Learning-as-a-Service (MLaaS) has become a widespread paradigm, making even the most complex machine learning models available for clients via e.g. a pay-per-query principle. This allows users to avoid time-consuming processes of data collection, hyperparameter tuning, and model training. However, by giving their customers access to the (predictions of their) models, MLaaS providers endanger their intellectual property, such as sensitive training data, optimised hyperparameters, or learned model parameters. Adversaries can create a copy of the model with (almost) identical behavior using the the prediction labels only. While many variants of this attack have been described, only scattered defence strategies have been proposed, addressing isolated threats. This raises the necessity for a thorough systematisation of the field of model stealing, to arrive at a comprehensive understanding why these attacks are successful, and how they could be holistically defended against. We address this by categorising and comparing model stealing attacks, assessing their performance, and exploring corresponding defence techniques in different settings. We propose a taxonomy for attack and defence approaches, and provide guidelines on how to select the right attack or defence strategy based on the goal and available resources. Finally, we analyse which defences are rendered less effective by current attack strategies.
Accepted at ACM Computing Surveys, 2023: https://doi.org/10.1145/3595292
References in corpus (11)
- Robust Machine Learning Systems: Challenges, Current Trends, Perspectives, and the Road Ahead
- Identifying Appropriate Intellectual Property Protection Mechanisms for Machine Learning Models: A Systematization of Watermarking, Fingerprinting, Model Access, and Attacks
- MEGEX: Data-Free Model Extraction Attack against Gradient-Based Explainable AI
- Stateful Detection of Model Extraction Attacks
- Thief, Beware of What Get You There: Towards Understanding Model Extraction Attack
- Perturbing Inputs to Prevent Model Stealing
- DynaMarks: Defending Against Deep Learning Model Extraction Using Dynamic Watermarking
- Good Artists Copy, Great Artists Steal: Model Extraction Attacks Against Image Translation Models
- Demystifying Arch-hints for Model Extraction: An Attack in Unified Memory System
- HODA: Hardness-Oriented Detection of Model Extraction Attacks
- Can't Steal? Cont-Steal! Contrastive Stealing Attacks Against Image Encoders
Cited by in corpus (10)
- Unleashing the potential of prompt engineering for large language models
- The Federation Strikes Back: A Survey of Federated Learning Privacy Attacks, Defenses, Applications, and Policy Landscape
- Identifying Appropriate Intellectual Property Protection Mechanisms for Machine Learning Models: A Systematization of Watermarking, Fingerprinting, Model Access, and Attacks
- An Intelligent Native Network Slicing Security Architecture Empowered by Federated Learning
- A PUF-Based Approach for Copy Protection of Intellectual Property in Neural Network Models
- Stealing the Invisible: Unveiling Pre-Trained CNN Models through Adversarial Examples and Timing Side-Channels
- Clone What You Can't Steal: Black-Box LLM Replication via Logit Leakage and Distillation
- Model Privacy: A Unified Framework for Understanding Model Stealing Attacks and Defenses
- Defending against Model Extraction for GNNs with Model Reprogramming
- ArcGen: Generalizing Neural Backdoor Detection Across Diverse Architectures