Februus: Input Purification Defense Against Trojan Attacks on Deep Neural Network Systems
arXiv:1908.03369 · doi:10.1145/3427228.3427264
Abstract
We propose Februus; a new idea to neutralize highly potent and insidious Trojan attacks on Deep Neural Network (DNN) systems at run-time. In Trojan attacks, an adversary activates a backdoor crafted in a deep neural network model using a secret trigger, a Trojan, applied to any input to alter the model's decision to a target prediction---a target determined by and only known to the attacker. Februus sanitizes the incoming input by surgically removing the potential trigger artifacts and restoring the input for the classification task. Februus enables effective Trojan mitigation by sanitizing inputs with no loss of performance for sanitized inputs, Trojaned or benign. Our extensive evaluations on multiple infected models based on four popular datasets across three contrasting vision applications and trigger types demonstrate the high efficacy of Februus. We dramatically reduced attack success rates from 100% to near 0% for all cases (achieving 0% on multiple cases) and evaluated the generalizability of Februus to defend against complex adaptive attacks; notably, we realized the first defense against the advanced partial Trojan attack. To the best of our knowledge, Februus is the first backdoor defense method for operation at run-time capable of sanitizing Trojaned inputs without requiring anomaly detection methods, model retraining or costly labeled data.
16 pages, to appear in the 36th Annual Computer Security Applications Conference (ACSAC 2020)
References in corpus (10)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Explaining and Harnessing Adversarial Examples
- Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning
- DeepID3: Face Recognition with Very Deep Neural Networks
- TABOR: A Highly Accurate Approach to Inspecting and Restoring Trojan Backdoors in AI Systems
- Blind Backdoors in Deep Learning Models
- On Certifying Robustness against Backdoor Attacks via Randomized Smoothing
- RAB: Provable Robustness Against Backdoor Attacks
- Design and Evaluation of a Multi-Domain Trojan Detection Method on Deep Neural Networks
- Backdoor Attacks to Graph Neural Networks
Cited by in corpus (35)
- AI Agents Under Threat: A Survey of Key Security Challenges and Future Pathways
- Input-Aware Dynamic Backdoor Attack
- Wild Patterns Reloaded: A Survey of Machine Learning Security against Training Data Poisoning
- Backdoor Attacks and Countermeasures on Deep Learning: A Comprehensive Review
- WaNet -- Imperceptible Warping-based Backdoor Attack
- Poison Ink: Robust and Invisible Backdoor Attack
- Rethinking the Trigger of Backdoor Attack
- AI Security for Geoscience and Remote Sensing: Challenges and Future Trends
- Backdoor Pre-trained Models Can Transfer to All
- You Autocomplete Me: Poisoning Vulnerabilities in Neural Code Completion
- Transferable Graph Backdoor Attack
- Reflection Backdoor: A Natural Backdoor Attack on Deep Neural Networks
- Spinning Language Models: Risks of Propaganda-As-A-Service and Countermeasures
- What Do You See? Evaluation of Explainable Artificial Intelligence (XAI) Interpretability through Neural Backdoors
- BadCM: Invisible Backdoor Attack Against Cross-Modal Learning
- Backdoor Attack through Frequency Domain
- Design and Evaluation of a Multi-Domain Trojan Detection Method on Deep Neural Networks
- FeSHI: Feature Map Based Stealthy Hardware Intrinsic Attack
- HaS-Nets: A Heal and Select Mechanism to Defend DNNs Against Backdoor Attacks for Data Collection Scenarios
- Light Can Hack Your Face! Black-box Backdoor Attack on Face Recognition Systems
- Two Souls in an Adversarial Image: Towards Universal Adversarial Example Detection using Multi-view Inconsistency
- DBIA: Data-free Backdoor Injection Attack against Transformer Networks
- NBA: defensive distillation for backdoor removal via neural behavior alignment
- NTD: Non-Transferability Enabled Backdoor Detection
- Two Sides of the Same Coin: Learning the Backdoor to Remove the Backdoor
- A General Framework for Defending Against Backdoor Attacks via Influence Graph
- DLP: towards active defense against backdoor attacks with decoupled learning process
- Towards Practical Deployment-Stage Backdoor Attack on Deep Neural Networks
- Live Trojan Attacks on Deep Neural Networks
- RAP: Robustness-Aware Perturbations for Defending against Backdoor Attacks on NLP Models
- Spinning Sequence-to-Sequence Models with Meta-Backdoors
- SGBA: A Stealthy Scapegoat Backdoor Attack against Deep Neural Networks
- The Victim and The Beneficiary: Exploiting a Poisoned Model to Train a Clean Model on Poisoned Data
- SpeechGuard: Online Defense against Backdoor Attacks on Speech Recognition Models
- Hypnopaedia-Aware Machine Unlearning via Psychometrics of Artificial Mental Imagery