High Accuracy and High Fidelity Extraction of Neural Networks
arXiv:1909.01838
Abstract
In a model extraction attack, an adversary steals a copy of a remotely deployed machine learning model, given oracle prediction access. We taxonomize model extraction attacks around two objectives: *accuracy*, i.e., performing well on the underlying learning task, and *fidelity*, i.e., matching the predictions of the remote victim classifier on any input. To extract a high-accuracy model, we develop a learning-based attack exploiting the victim to supervise the training of an extracted model. Through analytical and empirical arguments, we then explain the inherent limitations that prevent any learning-based strategy from extracting a truly high-fidelity model---i.e., extracting a functionally-equivalent model whose predictions are identical to those of the victim model on all possible inputs. Addressing these limitations, we expand on prior work to develop the first practical functionally-equivalent extraction attack for direct extraction (i.e., without training) of a model's weights. We perform experiments both on academic datasets and a state-of-the-art image classifier trained with 1 billion proprietary images. In addition to broadening the scope of model extraction research, our work demonstrates the practicality of model extraction attacks against production-grade systems.
USENIX Security 2020, 18 pages, 6 figures
References in corpus (7)
- Distilling the Knowledge in a Neural Network
- Sequence to Sequence Learning with Neural Networks
- Fashion-MNIST: a Novel Image Dataset for Benchmarking Machine Learning Algorithms
- WaveNet: A Generative Model for Raw Audio
- A framework for the extraction of Deep Neural Networks by leveraging public data
- Combining MixMatch and Active Learning for Better Accuracy with Fewer Labels
- On the Learnability of Deep Random Networks
Cited by in corpus (18)
- When Machine Unlearning Jeopardizes Privacy
- Thieves on Sesame Street! Model Extraction of BERT-based APIs
- Entangled Watermarks as a Defense against Model Extraction
- Mind Your Weight(s): A Large-scale Study on Insufficient Machine Learning Model Protection in Mobile Apps
- GAMIN: An Adversarial Approach to Black-Box Model Inversion
- Deep Neural Network Fingerprinting by Conferrable Adversarial Examples
- Fusion: Efficient and Secure Inference Resilient to Malicious Servers
- Reverse-Engineering Deep ReLU Networks
- Sneaky Spikes: Uncovering Stealthy Backdoor Attacks in Spiking Neural Networks with Neuromorphic Data
- Adversarial Model Extraction on Graph Neural Networks
- Improving LIME Robustness with Smarter Locality Sampling
- Thief, Beware of What Get You There: Towards Understanding Model Extraction Attack
- Extraction of Complex DNN Models: Real Threat or Boogeyman?
- BoMaNet: Boolean Masking of an Entire Neural Network
- Perturbing Inputs to Prevent Model Stealing
- ShadowNet: A Secure and Efficient On-device Model Inference System for Convolutional Neural Networks
- Theoretical Guarantees for Model Auditing with Finite Adversaries
- Power-Based Attacks on Spatial DNN Accelerators