BadNets: Identifying Vulnerabilities in the Machine Learning Model Supply Chain
arXiv:1708.06733
Abstract
Deep learning-based techniques have achieved state-of-the-art performance on a wide variety of recognition and classification tasks. However, these networks are typically computationally expensive to train, requiring weeks of computation on many GPUs; as a result, many users outsource the training procedure to the cloud or rely on pre-trained models that are then fine-tuned for a specific task. In this paper we show that outsourced training introduces new security risks: an adversary can create a maliciously trained network (a backdoored neural network, or a \emph{BadNet}) that has state-of-the-art performance on the user's training and validation samples, but behaves badly on specific attacker-chosen inputs. We first explore the properties of BadNets in a toy example, by creating a backdoored handwritten digit classifier. Next, we demonstrate backdoors in a more realistic scenario by creating a U.S. street sign classifier that identifies stop signs as speed limits when a special sticker is added to the stop sign; we then show in addition that the backdoor in our US street sign detector can persist even if the network is later retrained for another task and cause a drop in accuracy of {25}\% on average when the backdoor trigger is present. These results demonstrate that backdoors in neural networks are both powerful and---because the behavior of neural networks is difficult to explicate---stealthy. This work provides motivation for further research into techniques for verifying and inspecting neural networks, just as we have developed tools for verifying and debugging software.
References in corpus (4)
Cited by in corpus (244)
- Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning
- How To Backdoor Federated Learning
- Analyzing Federated Learning through an Adversarial Lens
- Spectral Signatures in Backdoor Attacks
- Mitigating Sybils in Federated Learning Poisoning
- Unleashing the potential of prompt engineering for large language models
- Februus: Input Purification Defense Against Trojan Attacks on Deep Neural Network Systems
- Threats to Federated Learning: A Survey
- Detecting Backdoor Attacks on Deep Neural Networks by Activation Clustering
- Ditto: Fair and Robust Federated Learning Through Personalization
- BadNL: Backdoor Attacks against NLP Models with Semantic-preserving Improvements
- Adversarial Example Detection for DNN Models: A Review and Experimental Comparison
- When Machine Unlearning Jeopardizes Privacy
- Local Model Poisoning Attacks to Byzantine-Robust Federated Learning
- Adversarial attacks and defenses in explainable artificial intelligence: A survey
- Threat of Adversarial Attacks on Deep Learning in Computer Vision: A Survey
- Input-Aware Dynamic Backdoor Attack
- Neural Attention Distillation: Erasing Backdoor Triggers from Deep Neural Networks
- TABOR: A Highly Accurate Approach to Inspecting and Restoring Trojan Backdoors in AI Systems
- Backdoor Attacks and Countermeasures on Deep Learning: A Comprehensive Review
- WaNet -- Imperceptible Warping-based Backdoor Attack
- Fooling Neural Network Interpretations via Adversarial Model Manipulation
- Manipulating Machine Learning: Poisoning Attacks and Countermeasures for Regression Learning
- Blind Backdoors in Deep Learning Models
- Attack of the Tails: Yes, You Really Can Backdoor Federated Learning
- Backdoor Embedding in Convolutional Neural Network Models via Invisible Perturbation
- Shallow-Deep Networks: Understanding and Mitigating Network Overthinking
- When Machine Learning Meets Privacy: A Survey and Outlook
- Rethinking the Trigger of Backdoor Attack
- Why Do Adversarial Attacks Transfer? Explaining Transferability of Evasion and Poisoning Attacks
- Physical Adversarial Attack meets Computer Vision: A Decade Survey
- Unsolved Problems in ML Safety
- STRIP: A Defence Against Trojan Attacks on Deep Neural Networks
- Assuring the Machine Learning Lifecycle: Desiderata, Methods, and Challenges
- Towards Adversarial Malware Detection: Lessons Learned from PDF-based Attacks
- AI Security for Geoscience and Remote Sensing: Challenges and Future Trends
- Privacy-Preserving Machine Learning: Methods, Challenges and Directions
- Backdoor Pre-trained Models Can Transfer to All
- Gotta Catch 'Em All: Using Honeypots to Catch Adversarial Attacks on Neural Networks
- Demon in the Variant: Statistical Analysis of DNNs for Robust Backdoor Contamination Detection
- Auditing Differentially Private Machine Learning: How Private is Private SGD?
- Robust Anomaly Detection and Backdoor Attack Detection Via Differential Privacy
- NeuronInspect: Detecting Backdoors in Neural Networks via Output Explanations
- Federated Unlearning: A Survey on Methods, Design Guidelines, and Evaluation Metrics
- You Autocomplete Me: Poisoning Vulnerabilities in Neural Code Completion
- Mitigating Backdoor Attacks in Federated Learning
- On Certifying Robustness against Backdoor Attacks via Randomized Smoothing
- Security and Privacy Issues in Deep Learning
- DAWN: Dynamic Adversarial Watermarking of Neural Networks
- Just How Toxic is Data Poisoning? A Unified Benchmark for Backdoor and Data Poisoning Attacks
- Explanation-Guided Backdoor Poisoning Attacks Against Malware Classifiers
- PoTrojan: powerful neural-level trojan designs in deep learning models
- ModelDiff: Testing-Based DNN Similarity Comparison for Model Reuse Detection
- A Survey on Vulnerability of Federated Learning: A Learning Algorithm Perspective
- Adversarial Machine Learning -- Industry Perspectives
- Weight Poisoning Attacks on Pre-trained Models
- Backdoor Scanning for Deep Neural Networks through K-Arm Optimization
- RAB: Provable Robustness Against Backdoor Attacks
- Poisoning Attacks with Generative Adversarial Nets
- Poisoning Deep Learning Based Recommender Model in Federated Learning Scenarios
- Prompt as Triggers for Backdoor Attack: Examining the Vulnerability in Language Models
- Entangled Watermarks as a Defense against Model Extraction
- Invisible Backdoor Attacks on Deep Neural Networks via Steganography and Regularization
- Label-Consistent Backdoor Attacks
- SAFELearning: Enable Backdoor Detectability In Federated Learning With Secure Aggregation
- Red Alarm for Pre-trained Models: Universal Vulnerability to Neuron-Level Backdoor Attacks
- Have You Stolen My Model? Evasion Attacks Against Deep Neural Network Watermarking Techniques
- Reflection Backdoor: A Natural Backdoor Attack on Deep Neural Networks
- Biometric Backdoors: A Poisoning Attack Against Unsupervised Template Updating
- PatchGuard: A Provably Robust Defense against Adversarial Patches via Small Receptive Fields and Masking
- Adversarial Unlearning of Backdoors via Implicit Hypergradient
- Defending against Backdoor Attack on Deep Neural Networks
- On Adversarial Bias and the Robustness of Fair Machine Learning
- Backdoor attacks and defenses in feature-partitioned collaborative learning
- Living-Off-The-Land Command Detection Using Active Learning
- Backdoor Attack in the Physical World
- FLTrust: Byzantine-robust Federated Learning via Trust Bootstrapping
- Fine-Pruning: Defending Against Backdooring Attacks on Deep Neural Networks
- Machine Learning Security against Data Poisoning: Are We There Yet?
- BadPre: Task-agnostic Backdoor Attacks to Pre-trained NLP Foundation Models
- Bugs in our Pockets: The Risks of Client-Side Scanning
- Security for Machine Learning-based Systems: Attacks and Challenges during Training and Inference
- A Master Key Backdoor for Universal Impersonation Attack against DNN-based Face Verification
- SLAP: Improving Physical Adversarial Examples with Short-Lived Adversarial Perturbations
- Training-free Lexical Backdoor Attacks on Language Models
- What Do You See? Evaluation of Explainable Artificial Intelligence (XAI) Interpretability through Neural Backdoors
- Universal Detection of Backdoor Attacks via Density-based Clustering and Centroids Analysis
- BadCM: Invisible Backdoor Attack Against Cross-Modal Learning
- Adversarial Neural Network Inversion via Auxiliary Knowledge Alignment
- DeepSweep: An Evaluation Framework for Mitigating DNN Backdoor Attacks using Data Augmentation
- The Threat of Adversarial Attacks on Machine Learning in Network Security -- A Survey
- TrISec: Training Data-Unaware Imperceptible Security Attacks on Deep Neural Networks
- Effectiveness of Distillation Attack and Countermeasure on Neural Network Watermarking
- TrojDRL: Trojan Attacks on Deep Reinforcement Learning Agents
- ConFoc: Content-Focus Protection Against Trojan Attacks on Neural Networks
- Neural network fragile watermarking with no model performance degradation
- Deep Neural Network Fingerprinting by Conferrable Adversarial Examples
- Poisoning the Unlabeled Dataset of Semi-Supervised Learning
- A Targeted Attack on Black-Box Neural Machine Translation with Parallel Data Poisoning
- Practical Detection of Trojan Neural Networks: Data-Limited and Data-Free Cases
- The TrojAI Software Framework: An OpenSource tool for Embedding Trojans into Deep Learning Models
- BAAAN: Backdoor Attacks Against Autoencoder and GAN-Based Machine Learning Models
- TrojanZoo: Towards Unified, Holistic, and Practical Evaluation of Neural Backdoors
- Backdoor Attacks on Federated Meta-Learning
- Handcrafted Backdoors in Deep Neural Networks
- Backdoor Attack through Frequency Domain
- Backdoor Attacks for Remote Sensing Data with Wavelet Transform
- DeepView: Visualizing Classification Boundaries of Deep Neural Networks as Scatter Plots Using Discriminative Dimensionality Reduction
- Turning Federated Learning Systems Into Covert Channels
- Design of intentional backdoors in sequential models
- Undistillable: Making A Nasty Teacher That CANNOT teach students
- A backdoor attack against LSTM-based text classification systems
- Poisoning and Backdooring Contrastive Learning
- PatchCleanser: Certifiably Robust Defense against Adversarial Patches for Any Image Classifier
- Evaluation of Inference Attack Models for Deep Learning on Medical Data
- Deep-Dup: An Adversarial Weight Duplication Attack Framework to Crush Deep Neural Network in Multi-Tenant FPGA
- A framework for fostering transparency in shared artificial intelligence models by increasing visibility of contributions
- Don't Trigger Me! A Triggerless Backdoor Attack Against Deep Neural Networks
- Clean-Label Backdoor Attacks on Video Recognition Models
- DP-InstaHide: Provably Defusing Poisoning and Backdoor Attacks with Differentially Private Data Augmentations
- Removing Backdoor-Based Watermarks in Neural Networks with Limited Data
- Poison as a Cure: Detecting & Neutralizing Variable-Sized Backdoor Attacks in Deep Neural Networks
- AEVA: Black-box Backdoor Detection Using Adversarial Extreme Value Analysis
- Review and Critical Analysis of Privacy-preserving Infection Tracking and Contact Tracing
- Shielding Collaborative Learning: Mitigating Poisoning Attacks through Client-Side Detection
- Walling up Backdoors in Intrusion Detection Systems
- Programmable Neural Network Trojan for Pre-Trained Feature Extractor
- Design and Evaluation of a Multi-Domain Trojan Detection Method on Deep Neural Networks
- Membership Leakage in Label-Only Exposures
- Rethinking the Backdoor Attacks' Triggers: A Frequency Perspective
- How to Manipulate CNNs to Make Them Lie: the GradCAM Case
- Poisoned classifiers are not only backdoored, they are fundamentally broken
- Learning to Confuse: Generating Training Time Adversarial Data with Auto-Encoder
- RED-Attack: Resource Efficient Decision based Attack for Machine Learning
- Light Can Hack Your Face! Black-box Backdoor Attack on Face Recognition Systems
- Orchestrating the Development Lifecycle of Machine Learning-Based IoT Applications: A Taxonomy and Survey
- Geometry-Inspired Top-k Adversarial Perturbations
- Security and Privacy for Artificial Intelligence: Opportunities and Challenges
- T-Miner: A Generative Approach to Defend Against Trojan Attacks on DNN-based Text Classification
- VerIDeep: Verifying Integrity of Deep Neural Networks through Sensitive-Sample Fingerprinting
- Hidden Killer: Invisible Textual Backdoor Attacks with Syntactic Trigger
- TBT: Targeted Neural Network Attack with Bit Trojan
- Backdoor Attacks to Graph Neural Networks
- SPECTRE: Defending Against Backdoor Attacks Using Robust Statistics
- An Embarrassingly Simple Approach for Trojan Attack in Deep Neural Networks
- Embedding and Extraction of Knowledge in Tree Ensemble Classifiers
- LMSanitator: Defending Prompt-Tuning Against Task-Agnostic Backdoors
- TOP: Backdoor Detection in Neural Networks via Transferability of Perturbation
- Poisoning Deep Reinforcement Learning Agents with In-Distribution Triggers
- Poisoning Programs by Un-Repairing Code: Security Concerns of AI-generated Code
- Be Careful about Poisoned Word Embeddings: Exploring the Vulnerability of the Embedding Layers in NLP Models
- FooBaR: Fault Fooling Backdoor Attack on Neural Network Training
- SoK: Machine Learning Governance
- M-to-N Backdoor Paradigm: A Multi-Trigger and Multi-Target Attack to Deep Learning Models
- Backdoor Attacks on Crowd Counting
- The Feasibility and Inevitability of Stealth Attacks
- ImpNet: Imperceptible and blackbox-undetectable backdoors in compiled neural networks
- One-pixel Signature: Characterizing CNN Models for Backdoor Detection
- Poisoning MorphNet for Clean-Label Backdoor Attack to Point Clouds
- Lower Bounds for Adversarially Robust PAC Learning
- Poisoning Attacks to Local Differential Privacy Protocols for Key-Value Data
- Defending Against Backdoor Attacks in Natural Language Generation
- Sleeper Agent: Scalable Hidden Trigger Backdoors for Neural Networks Trained from Scratch
- Deep Serial Number: Computational Watermarking for DNN Intellectual Property Protection
- Turn the Combination Lock: Learnable Textual Backdoor Attacks via Word Substitution
- What Do Deep Nets Learn? Class-wise Patterns Revealed in the Input Space
- TAD: Trigger Approximation based Black-box Trojan Detection for AI
- Membership Inference Attacks Against Recommender Systems
- PRECAD: Privacy-Preserving and Robust Federated Learning via Crypto-Aided Differential Privacy
- Manipulating SGD with Data Ordering Attacks
- GOAT: GPU Outsourcing of Deep Learning Training With Asynchronous Probabilistic Integrity Verification Inside Trusted Execution Environment
- FaceHack: Triggering backdoored facial recognition systems using facial characteristics
- BACKDOORL: Backdoor Attack against Competitive Reinforcement Learning
- DBIA: Data-free Backdoor Injection Attack against Transformer Networks
- Adversarial Attacks and Defenses: An Interpretation Perspective
- Regula Sub-rosa: Latent Backdoor Attacks on Deep Neural Networks
- Backdoor Attacks on Federated Learning with Lottery Ticket Hypothesis
- FIBA: Frequency-Injection based Backdoor Attack in Medical Image Analysis
- Benchmarking Spiking Neural Network Learning Methods with Varying Locality
- HoneyModels: Machine Learning Honeypots
- FedCom: A Byzantine-Robust Local Model Aggregation Rule Using Data Commitment for Federated Learning
- Built-in Vulnerabilities to Imperceptible Adversarial Perturbations
- Influence Function based Data Poisoning Attacks to Top-N Recommender Systems
- Poison Attacks against Text Datasets with Conditional Adversarially Regularized Autoencoder
- Detecting Localized Adversarial Examples: A Generic Approach using Critical Region Analysis
- Adaptive Backdoor Attacks with Reasonable Constraints on Graph Neural Networks
- Gradient Shaping: Enhancing Backdoor Attack Against Reverse Engineering
- Thundernna: a white box adversarial attack
- HufuNet: Embedding the Left Piece as Watermark and Keeping the Right Piece for Ownership Verification in Deep Neural Networks
- NTD: Non-Transferability Enabled Backdoor Detection
- A Unified Framework for Task-Driven Data Quality Management
- DeepPoison: Feature Transfer Based Stealthy Poisoning Attack
- Qu-ANTI-zation: Exploiting Quantization Artifacts for Achieving Adversarial Outcomes
- Poisoned Source Code Detection in Code Models
- Two Sides of the Same Coin: Learning the Backdoor to Remove the Backdoor
- Good-looking but Lacking Faithfulness: Understanding Local Explanation Methods through Trend-based Testing
- Amplification trojan network: Attack deep neural networks by amplifying their inherent weakness
- Active Learning Under Malicious Mislabeling and Poisoning Attacks
- ONION: A Simple and Effective Defense Against Textual Backdoor Attacks
- Machine Learning with Electronic Health Records is vulnerable to Backdoor Trigger Attacks
- Mind the Style of Text! Adversarial and Backdoor Attacks Based on Text Style Transfer
- TESDA: Transform Enabled Statistical Detection of Attacks in Deep Neural Networks
- Towards Compliant Data Management Systems for Healthcare ML
- DeepPayload: Black-box Backdoor Attack on Deep Learning Models through Neural Payload Injection
- A General Framework for Defending Against Backdoor Attacks via Influence Graph
- VerifyTL: Secure and Verifiable Collaborative Transfer Learning
- Decamouflage: A Framework to Detect Image-Scaling Attacks on Convolutional Neural Networks
- Topological Detection of Trojaned Neural Networks
- Incorrect by Construction: Fine Tuning Neural Networks for Guaranteed Performance on Finite Sets of Examples
- Reaching Data Confidentiality and Model Accountability on the CalTrain
- Quantifying Transparency of Machine Learning Systems through Analysis of Contributions
- Lightweight machine unlearning in neural network
- Live Trojan Attacks on Deep Neural Networks
- Robust Risk Minimization for Statistical Learning
- PointBA: Towards Backdoor Attacks in 3D Point Cloud
- On Provable Backdoor Defense in Collaborative Learning
- Being Single Has Benefits. Instance Poisoning to Deceive Malware Classifiers
- Using Randomness to Improve Robustness of Machine-Learning Models Against Evasion Attacks
- Putting words into the system's mouth: A targeted attack on neural machine translation using monolingual data poisoning
- Towards Scheduling Federated Deep Learning using Meta-Gradients for Inter-Hospital Learning
- Who is Responsible for Adversarial Defense?
- Technical Report: When Does Machine Learning FAIL? Generalized Transferability for Evasion and Poisoning Attacks
- Semantically Robust Unpaired Image Translation for Data with Unmatched Semantics Statistics
- Excess Capacity and Backdoor Poisoning
- 10 Security and Privacy Problems in Large Foundation Models
- SGBA: A Stealthy Scapegoat Backdoor Attack against Deep Neural Networks
- Learning and Certification under Instance-targeted Poisoning
- Incompatibility Clustering as a Defense Against Backdoor Poisoning Attacks
- De-Pois: An Attack-Agnostic Defense against Data Poisoning Attacks
- Spinning Sequence-to-Sequence Models with Meta-Backdoors
- The Victim and The Beneficiary: Exploiting a Poisoned Model to Train a Clean Model on Poisoned Data
- RAP: Robustness-Aware Perturbations for Defending against Backdoor Attacks on NLP Models
- Risk Management Framework for Machine Learning Security
- Backdoor Attacks on Pre-trained Models by Layerwise Weight Poisoning
- On the Security Risks of AutoML
- A Separation Result Between Data-oblivious and Data-aware Poisoning Attacks
- i-Algebra: Towards Interactive Interpretability of Deep Neural Networks
- TRAPDOOR: Repurposing backdoors to detect dataset bias in machine learning-based genomic analysis
- Adversarial examples are useful too!
- Law and Adversarial Machine Learning
- Accumulative Poisoning Attacks on Real-time Data
- High-Robustness, Low-Transferability Fingerprinting of Neural Networks
- Semantic Host-free Trojan Attack
- PoisHygiene: Detecting and Mitigating Poisoning Attacks in Neural Networks