Practical Black-Box Attacks against Machine Learning
arXiv:1602.02697
Abstract
Machine learning (ML) models, e.g., deep neural networks (DNNs), are vulnerable to adversarial examples: malicious inputs modified to yield erroneous model outputs, while appearing unmodified to human observers. Potential attacks include having malicious content like malware identified as legitimate or controlling vehicle behavior. Yet, all existing adversarial example attacks require knowledge of either the model internals or its training data. We introduce the first practical demonstration of an attacker controlling a remotely hosted DNN with no such knowledge. Indeed, the only capability of our black-box adversary is to observe labels given by the DNN to chosen inputs. Our attack strategy consists in training a local model to substitute for the target DNN, using inputs synthetically generated by an adversary and labeled by the target DNN. We use the local substitute to craft adversarial examples, and find that they are misclassified by the targeted DNN. To perform a real-world and properly-blinded evaluation, we attack a DNN hosted by MetaMind, an online deep learning API. We find that their DNN misclassifies 84.24% of the adversarial examples crafted with our substitute. We demonstrate the general applicability of our strategy to many ML techniques by conducting the same attack against models hosted by Amazon and Google, using logistic regression substitutes. They yield adversarial examples misclassified by Amazon and Google at rates of 96.19% and 88.94%. We also find that this black-box attack strategy is capable of evading defense strategies previously found to make adversarial example crafting harder.
Proceedings of the 2017 ACM Asia Conference on Computer and Communications Security, Abu Dhabi, UAE
References in corpus (6)
- Distilling the Knowledge in a Neural Network
- Theano: new features and speed improvements
- Poisoning Attacks against Support Vector Machines
- Stealing Machine Learning Models via Prediction APIs
- Towards Deep Neural Network Architectures Robust to Adversarial Examples
- Evasion and Hardening of Tree Ensemble Classifiers
Cited by in corpus (72)
- Feature Squeezing: Detecting Adversarial Examples in Deep Neural Networks
- Transferability in Machine Learning: from Phenomena to Black-Box Attacks using Adversarial Samples
- BadNets: Identifying Vulnerabilities in the Machine Learning Model Supply Chain
- The Space of Transferable Adversarial Examples
- Generating Adversarial Malware Examples for Black-Box Attacks Based on GAN
- On the (Statistical) Detection of Adversarial Examples
- Certifying Some Distributional Robustness with Principled Adversarial Training
- Adversarial Perturbations Against Deep Neural Networks for Malware Classification
- AdvHat: Real-world adversarial attack on ArcFace Face ID system
- Identifying Implementation Bugs in Machine Learning based Image Classifiers using Metamorphic Testing
- Houdini: Fooling Deep Structured Prediction Models
- Adversarial Attack on Graph Structured Data
- Predicting the Generalization Gap in Deep Networks with Margin Distributions
- The Robust Manifold Defense: Adversarial Training using Generative Models
- Deceiving End-to-End Deep Learning Malware Detectors using Adversarial Examples
- Understanding and Enhancing the Transferability of Adversarial Examples
- Black-Box Attacks against RNN based Malware Detection Algorithms
- Adversarial Examples for Semantic Segmentation and Object Detection
- Blocking Transferability of Adversarial Examples in Black-Box Learning Systems
- APE-GAN: Adversarial Perturbation Elimination with GAN
- Adversarial Example Defenses: Ensembles of Weak Defenses are not Strong
- Ensemble Methods as a Defense to Adversarial Perturbations Against Deep Neural Networks
- Boosting Adversarial Training with Hypersphere Embedding
- Black-box Generation of Adversarial Text Sequences to Evade Deep Learning Classifiers
- Boosting Adversarial Attacks with Momentum
- Towards Robust Neural Networks via Random Self-ensemble
- SafetyNet: Detecting and Rejecting Adversarial Examples Robustly
- Attributing Fake Images to GANs: Learning and Analyzing GAN Fingerprints
- Whatever Does Not Kill Deep Reinforcement Learning, Makes It Stronger
- Machine Learning (In) Security: A Stream of Problems
- Adversarial examples for generative models
- Black-box Adversarial Attacks with Bayesian Optimization
- Intriguing Properties of Adversarial Examples
- Benchmarking Adversarial Robustness
- Assessing Threat of Adversarial Examples on Deep Neural Networks
- Rethinking Non-idealities in Memristive Crossbars for Adversarial Robustness in Neural Networks
- Regional Homogeneity: Towards Learning Transferable Universal Adversarial Perturbations Against Defenses
- Percival: Making In-Browser Perceptual Ad Blocking Practical With Deep Learning
- Security and Privacy for Artificial Intelligence: Opportunities and Challenges
- Noise Sensitivity-Based Energy Efficient and Robust Adversary Detection in Neural Networks
- Towards Robust Detection of Adversarial Examples
- Transferability of Adversarial Examples to Attack Cloud-based Image Classifier Service
- Google's Cloud Vision API Is Not Robust To Noise
- Adversarial Image Translation: Unrestricted Adversarial Examples in Face Recognition Systems
- Fooling Vision and Language Models Despite Localization and Attention Mechanism
- Building Robust Deep Neural Networks for Road Sign Detection
- How intelligent are convolutional neural networks?
- Attacking Automatic Video Analysis Algorithms: A Case Study of Google Cloud Video Intelligence API
- Evaluation of Momentum Diverse Input Iterative Fast Gradient Sign Method (M-DI2-FGSM) Based Attack Method on MCS 2018 Adversarial Attacks on Black Box Face Recognition System
- Patch-wise++ Perturbation for Adversarial Targeted Attacks
- A3T: Adversarially Augmented Adversarial Training
- Attacking and Defending Machine Learning Applications of Public Cloud
- Second Order Optimization for Adversarial Robustness and Interpretability
- Maximum Resilience of Artificial Neural Networks
- Fraternal Twins: Unifying Attacks on Machine Learning and Digital Watermarking
- Proof-of-Learning: Definitions and Practice
- Robustness Certificates Against Adversarial Examples for ReLU Networks
- A Comprehensive Evaluation Framework for Deep Model Robustness
- Models and Framework for Adversarial Attacks on Complex Adaptive Systems
- RAIN: A Simple Approach for Robust and Accurate Image Classification Networks
- Reinforcement Learning for Autonomous Defence in Software-Defined Networking
- Towards Deep Learning Models Resistant to Large Perturbations
- A New Angle on L2 Regularization
- Detecting Adversarial Samples Using Density Ratio Estimates
- The best defense is a good offense: Countering black box attacks by predicting slightly wrong labels
- Defending from adversarial examples with a two-stream architecture
- Simple and Efficient Hard Label Black-box Adversarial Attacks in Low Query Budget Regimes
- Metamorphic Testing of a Deep Learning based Forecaster
- Power-Based Attacks on Spatial DNN Accelerators
- Adversarial Attacks Against Deep Learning Systems for ICD-9 Code Assignment
- Learning To Characterize Adversarial Subspaces
- Image Decomposition and Classification through a Generative Model