Towards Reverse-Engineering Black-Box Neural Networks
arXiv:1711.01768
Abstract
Many deployed learned models are black boxes: given input, returns output. Internal information about the model, such as the architecture, optimisation procedure, or training data, is not disclosed explicitly as it might contain proprietary information or make the system more vulnerable. This work shows that such attributes of neural networks can be exposed from a sequence of queries. This has multiple implications. On the one hand, our work exposes the vulnerability of black-box neural networks to different types of attacks -- we show that the revealed internal information helps generate more effective adversarial examples against the black box model. On the other hand, this technique can be used for better protection of private content from automatic recognition models using adversarial examples. Our paper suggests that it is actually hard to draw a line between white box and black box models.
20 pages, 12 figures, to appear at ICLR'18. Code: https://goo.gl/MbYfsv
Cited by in corpus (18)
- A Survey of Privacy Attacks in Machine Learning
- Improving Black-box Adversarial Attacks with a Transfer-based Prior
- Privacy-preserving Machine Learning through Data Obfuscation
- Defending Model Inversion and Membership Inference Attacks via Prediction Purification
- A framework for the extraction of Deep Neural Networks by leveraging public data
- Attributing Fake Images to GANs: Learning and Analyzing GAN Fingerprints
- Black-Box Ripper: Copying black-box models using generative evolutionary algorithms
- IPGuard: Protecting Intellectual Property of Deep Neural Networks via Fingerprinting the Classification Boundary
- An Overview of Privacy in Machine Learning
- Segmentations-Leak: Membership Inference Attacks and Defenses in Semantic Image Segmentation
- Quantifying and Mitigating Privacy Risks of Contrastive Learning
- DeepPeep: Exploiting Design Ramifications to Decipher the Architecture of Compact DNNs
- SEALing Neural Network Models in Secure Deep Learning Accelerators
- Teacher Model Fingerprinting Attacks Against Transfer Learning
- BODAME: Bilevel Optimization for Defense Against Model Extraction
- Stealing Black-Box Functionality Using The Deep Neural Tree Architecture
- Mental Models of Adversarial Machine Learning
- A Novel Privacy-Preserving Deep Learning Scheme without Using Cryptography Component