DeepXplore: Automated Whitebox Testing of Deep Learning Systems
arXiv:1705.06640 · doi:10.1145/3132747.3132785
Abstract
Deep learning (DL) systems are increasingly deployed in safety- and security-critical domains including self-driving cars and malware detection, where the correctness and predictability of a system's behavior for corner case inputs are of great importance. Existing DL testing depends heavily on manually labeled data and therefore often fails to expose erroneous behaviors for rare inputs. We design, implement, and evaluate DeepXplore, the first whitebox framework for systematically testing real-world DL systems. First, we introduce neuron coverage for systematically measuring the parts of a DL system exercised by test inputs. Next, we leverage multiple DL systems with similar functionality as cross-referencing oracles to avoid manual checking. Finally, we demonstrate how finding inputs for DL systems that both trigger many differential behaviors and achieve high neuron coverage can be represented as a joint optimization problem and solved efficiently using gradient-based search techniques. DeepXplore efficiently finds thousands of incorrect corner case behaviors (e.g., self-driving cars crashing into guard rails and malware masquerading as benign software) in state-of-the-art DL models with thousands of neurons trained on five popular datasets including ImageNet and Udacity self-driving challenge data. For all tested DL models, on average, DeepXplore generated one test input demonstrating incorrect behavior within one second while running only on a commodity laptop. We further show that the test inputs generated by DeepXplore can also be used to retrain the corresponding DL model to improve the model's accuracy by up to 3%.
To be published in SOSP'17
References in corpus (7)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Understanding Neural Networks Through Deep Visualization
- Stealing Machine Learning Models via Prediction APIs
- Towards Deep Neural Network Architectures Robust to Adversarial Examples
- Learning to Generate Reviews and Discovering Sentiment
- Parseval Networks: Improving Robustness to Adversarial Examples
- Extending Defensive Distillation
Cited by in corpus (252)
- Wild Patterns: Ten Years After the Rise of Adversarial Machine Learning
- DeepGauge: Multi-Granularity Testing Criteria for Deep Learning Systems
- Generating Natural Adversarial Examples
- Guiding Deep Learning System Testing using Surprise Adequacy
- DLFuzz: Differential Fuzzing Testing of Deep Learning Systems
- A Survey of End-to-End Driving: Architectures and Training Methods
- Formal Security Analysis of Neural Networks using Symbolic Intervals
- Invisible for both Camera and LiDAR: Security of Multi-Sensor Fusion based Perception in Autonomous Driving Under Physical-World Attacks
- Adversarial Examples: Attacks and Defenses for Deep Learning
- DeepTest: Automated Testing of Deep-Neural-Network-driven Autonomous Cars
- Automated Directed Fairness Testing
- Identifying Implementation Bugs in Machine Learning based Image Classifiers using Metamorphic Testing
- Testing Deep Neural Networks
- A Performance-Sensitive Malware Detection System Using Deep Learning on Mobile Devices
- How Deep Learning Sees the World: A Survey on Adversarial Attacks & Defenses
- TensorFuzz: Debugging Neural Networks with Coverage-Guided Fuzzing
- TABOR: A Highly Accurate Approach to Inspecting and Restoring Trojan Backdoors in AI Systems
- Robust Machine Learning Systems: Challenges, Current Trends, Perspectives, and the Road Ahead
- Machine Learning Testing: Survey, Landscapes and Horizons
- Deep Learning Methods for Fingerprint-Based Indoor Positioning: A Review
- Mind the Gap! A Study on the Transferability of Virtual vs Physical-world Testing of Autonomous Driving Systems
- Failing to Learn: Autonomously Identifying Perception Failures for Self-driving Cars
- Do the Machine Learning Models on a Crowd Sourced Platform Exhibit Bias? An Empirical Study on Model Fairness
- Black-Box Testing of Deep Neural Networks Through Test Case Diversity
- How to Certify Machine Learning Based Safety-critical Systems? A Systematic Literature Review
- Verification for Machine Learning, Autonomy, and Neural Networks Survey
- Assuring the Machine Learning Lifecycle: Desiderata, Methods, and Challenges
- Optimization and Abstraction: A Synergistic Approach for Analyzing Neural Network Robustness
- Boosting Operational DNN Testing Efficiency through Conditioning
- DriveFuzz: Discovering Autonomous Driving Bugs through Driving Quality-Guided Fuzzing
- Towards Practical Verification of Machine Learning: The Case of Computer Vision Systems
- Arachne: Search Based Repair of Deep Neural Networks
- AVFI: Fault Injection for Autonomous Vehicles
- FDA3 : Federated Defense Against Adversarial Attacks for Cloud-Based IIoT Applications
- Robust Federated Learning: The Case of Affine Distribution Shifts
- WhiteFox: White-Box Compiler Fuzzing Empowered by Large Language Models
- DeepRoad: GAN-based Metamorphic Autonomous Driving System Testing
- Description of Corner Cases in Automated Driving: Goals and Challenges
- Combinatorial Testing for Deep Learning Systems
- Comparing Offline and Online Testing of Deep Neural Networks: An Autonomous Car Case Study
- TBD: Benchmarking and Analyzing Deep Neural Network Training
- A Survey of Safety and Trustworthiness of Deep Neural Networks: Verification, Testing, Adversarial Attack and Defence, and Interpretability
- Simple Techniques Work Surprisingly Well for Neural Network Test Prioritization and Active Learning (Replicability Study)
- A Safety Framework for Critical Systems Utilising Deep Neural Networks
- ReluDiff: Differential Verification of Deep Neural Networks
- ModelDiff: Testing-Based DNN Similarity Comparison for Model Reuse Detection
- Adversarial Explanations for Understanding Image Classification Decisions and Improved Neural Network Robustness
- Yes, we GAN: Applying Adversarial Techniques for Autonomous Driving
- Experimental Resilience Assessment of An Open-Source Driving Agent
- Guidance on the Assurance of Machine Learning in Autonomous Systems (AMLAS)
- Characterizing the Decision Boundary of Deep Neural Networks
- Reducing DNN Labelling Cost using Surprise Adequacy: An Industrial Case Study for Autonomous Driving
- Symbolic Execution for Deep Neural Networks
- Detecting Adversarial Samples for Deep Neural Networks through Mutation Testing
- Rigorous Agent Evaluation: An Adversarial Approach to Uncover Catastrophic Failures
- DeepBillboard: Systematic Physical-World Testing of Autonomous Driving Systems
- Causal Testing: Finding Defects' Root Causes
- Coverage Guided Testing for Recurrent Neural Networks
- Using Machine Learning Safely in Automotive Software: An Assessment and Adaption of Software Process Requirements in ISO 26262
- Fuzzing Automatic Differentiation in Deep-Learning Libraries
- Algorithms for Verifying Deep Neural Networks
- DeepGD: A Multi-Objective Black-Box Test Selection Approach for Deep Neural Networks
- Dynamic Slicing for Deep Neural Networks
- Understanding Performance Problems in Deep Learning Systems
- DeepCruiser: Automated Guided Testing for Stateful Deep Learning Systems
- On Misbehaviour and Fault Tolerance in Machine Learning Systems
- FakeSpotter: A Simple yet Robust Baseline for Spotting AI-Synthesized Fake Faces
- Testing DNN Image Classifiers for Confusion & Bias Errors
- Kayotee: A Fault Injection-based System to Assess the Safety and Reliability of Autonomous Vehicles to Faults and Errors
- Adversarial Deep Reinforcement Learning for Improving the Robustness of Multi-agent Autonomous Driving Policies
- DLA: Dense-Layer-Analysis for Adversarial Example Detection
- A Dataset for GitHub Repository Deduplication: Extended Description
- Paracosm: A Language and Tool for Testing Autonomous Driving Systems
- Verifying Controllers Against Adversarial Examples with Bayesian Optimization
- Towards a Robust Deep Neural Network in Texts: A Survey
- Increasing the Confidence of Deep Neural Networks by Coverage Analysis
- An Exploratory Study on the Introduction and Removal of Different Types of Technical Debt
- Exposing Previously Undetectable Faults in Deep Neural Networks
- Operational Calibration: Debugging Confidence Errors for DNNs in the Field
- Finding Critical Scenarios for Automated Driving Systems: A Systematic Literature Review
- Assessing test artifact quality -- A tertiary study
- A Survey on Verification and Validation, Testing and Evaluations of Neurosymbolic Artificial Intelligence
- DeepHunter: Hunting Deep Neural Network Defects via Coverage-Guided Fuzzing
- Dirty Road Can Attack: Security of Deep Learning based Automated Lane Centering under Physical-World Attack
- The RFML Ecosystem: A Look at the Unique Challenges of Applying Deep Learning to Radio Frequency Applications
- Automated Test Generation to Detect Individual Discrimination in AI Models
- Software Engineering Practice in the Development of Deep Learning Applications
- ConFoc: Content-Focus Protection Against Trojan Attacks on Neural Networks
- FairPrep: Promoting Data to a First-Class Citizen in Studies on Fairness-Enhancing Interventions
- Toward Certification of Machine-Learning Systems for Low Criticality Airborne Applications
- There is Limited Correlation between Coverage and Robustness for Deep Neural Networks
- Secure Deep Learning Engineering: A Software Quality Assurance Perspective
- DeepSonar: Towards Effective and Robust Detection of AI-Synthesized Fake Voices
- Metamorphic Testing for Object Detection Systems
- Model Assertions for Monitoring and Improving ML Models
- Performance Analysis of Out-of-Distribution Detection on Trained Neural Networks
- A Formalization of Robustness for Deep Neural Networks
- Exploring Adversarial Attack in Spiking Neural Networks with Spike-Compatible Gradient
- MultiTest: Physical-Aware Object Insertion for Testing Multi-sensor Fusion Perception Systems
- MULDEF: Multi-model-based Defense Against Adversarial Examples for Neural Networks
- Coverage Testing of Deep Learning Models using Dataset Characterization
- Patching Weak Convolutional Neural Network Models through Modularization and Composition
- Road Context-aware Intrusion Detection System for Autonomous Cars
- DeepObfuscation: Securing the Structure of Convolutional Neural Networks via Knowledge Distillation
- FedDebug: Systematic Debugging for Federated Learning Applications
- Metamorphic Relation Based Adversarial Attacks on Differentiable Neural Computer
- PatchCensor: Patch Robustness Certification for Transformers via Exhaustive Testing
- RAID: Randomized Adversarial-Input Detection for Neural Networks
- Automatic Techniques to Systematically Discover New Heap Exploitation Primitives
- TestRank: Bringing Order into Unlabeled Test Instances for Deep Learning Tasks
- FedDefender: Backdoor Attack Defense in Federated Learning
- MarMot: Metamorphic Runtime Monitoring of Autonomous Driving Systems
- PatchAttack: A Black-box Texture-based Attack with Reinforcement Learning
- An Empirical Study towards Characterizing Deep Learning Development and Deployment across Different Frameworks and Platforms
- Smoke Testing for Machine Learning: Simple Tests to Discover Severe Defects
- UnbiasedNets: A Dataset Diversification Framework for Robustness Bias Alleviation in Neural Networks
- Enhancing Gradient-based Attacks with Symbolic Intervals
- DeepCover: Advancing RNN Test Coverage and Online Error Prediction using State Machine Extraction
- DeepGini: Prioritizing Massive Tests to Enhance the Robustness of Deep Neural Networks
- A Marauder's Map of Security and Privacy in Machine Learning
- EREBA: Black-box Energy Testing of Adaptive Neural Networks
- Topological Differential Testing
- Correctness Verification of Neural Networks
- Lifecycle Management of Trustworthy AI Models in 6G Networks: The REASON Approach
- : Activation Anomaly Analysis
- Towards Logical Specification of Statistical Machine Learning
- Security and Machine Learning in the Real World
- Towards Testing of Deep Learning Systems with Training Set Reduction
- A Systematic Mapping Study on Testing of Machine Learning Programs
- Test Selection for Deep Learning Systems
- Wolf in Sheep's Clothing - The Downscaling Attack Against Deep Learning Applications
- Discovering Closed-Loop Failures of Vision-Based Controllers via Reachability Analysis
- CoCoFuzzing: Testing Neural Code Models with Coverage-Guided Fuzzing
- DiverGet: A Search-Based Software Testing Approach for Deep Neural Network Quantization Assessment
- Testing for Fault Diversity in Reinforcement Learning
- Beyond Accuracy: An Empirical Study on Unit Testing in Open-source Deep Learning Projects
- QuanTest: Entanglement-Guided Testing of Quantum Neural Network Systems
- DeltaNN: Assessing the Impact of Computational Environment Parameters on the Performance of Image Recognition Models
- ShapeFlow: Dynamic Shape Interpreter for TensorFlow
- Fairness Testing of Deep Image Classification with Adequacy Metrics
- Detecting Operational Adversarial Examples for Reliable Deep Learning
- Scalable Synthesis of Verified Controllers in Deep Reinforcement Learning
- Validation Frameworks for Self-Driving Vehicles: A Survey
- CAGFuzz: Coverage-Guided Adversarial Generative Fuzzing Testing of Deep Learning Systems
- Refactoring Neural Networks for Verification
- Visualizing and Understanding Deep Neural Networks in CTR Prediction
- Robust Semantic Segmentation with Superpixel-Mix
- An Epistemic Approach to the Formal Specification of Statistical Machine Learning
- ADI: Adversarial Dominating Inputs in Vertical Federated Learning Systems
- TSS: Transformation-Specific Smoothing for Robustness Certification
- Perfectly Parallel Fairness Certification of Neural Networks
- Opening the Software Engineering Toolbox for the Assessment of Trustworthy AI
- ART: Abstraction Refinement-Guided Training for Provably Correct Neural Networks
- A Comprehensive Evaluation Framework for Deep Model Robustness
- Revisiting Deep Neural Network Test Coverage from the Test Effectiveness Perspective
- Towards Characterizing Adversarial Defects of Deep Learning Software from the Lens of Uncertainty
- A Study of the Learnability of Relational Properties: Model Counting Meets Machine Learning (MCML)
- The Automated Inspection of Opaque Liquid Vaccines
- Quality attributes of test cases and test suites -- importance & challenges from practitioners' perspectives
- Deep Learning in the Automotive Industry: Recent Advances and Application Examples
- Rearchitecting Classification Frameworks For Increased Robustness
- An Empirical Study on Deployment Faults of Deep Learning Based Mobile Applications
- Data Sanity Check for Deep Learning Systems via Learnt Assertions
- Testing Deep Learning Models for Image Analysis Using Object-Relevant Metamorphic Relations
- testRNN: Coverage-guided Testing on Recurrent Neural Networks
- Mediation Challenges and Socio-Technical Gaps for Explainable Deep Learning Applications
- Metamorphic Testing of Image Captioning Systems via Image-Level Reduction
- NeuroDiff: Scalable Differential Verification of Neural Networks using Fine-Grained Approximation
- ArchRepair: Block-Level Architecture-Oriented Repairing for Deep Neural Networks
- On Machine Learning and Structure for Mobile Robots
- On Robustness of Lane Detection Models to Physical-World Adversarial Attacks in Autonomous Driving
- RobOT: Robustness-Oriented Testing for Deep Learning Systems
- Taking Care of The Discretization Problem: A Comprehensive Study of the Discretization Problem and A Black-Box Adversarial Attack in Discrete Integer Domain
- Synthesis-guided Adversarial Scenario Generation for Gray-box Feedback Control Systems with Sensing Imperfections
- Testing Untestable Neural Machine Translation: An Industrial Case
- PhysGAN: Generating Physical-World-Resilient Adversarial Examples for Autonomous Driving
- Testing Monotonicity of Machine Learning Models
- DeepSearch: A Simple and Effective Blackbox Attack for Deep Neural Networks
- TEASMA: A Practical Methodology for Test Adequacy Assessment of Deep Neural Networks
- Testing of Autonomous Driving Systems: Where Are We and Where Should We Go?
- Model-Based Robust Deep Learning: Generalizing to Natural, Out-of-Distribution Data
- Targeted Deep Learning System Boundary Testing
- Measuring Discrimination to Boost Comparative Testing for Multiple Deep Learning Models
- Feature-Filter: Detecting Adversarial Examples through Filtering off Recessive Features
- Can Offline Testing of Deep Neural Networks Replace Their Online Testing?
- Abstraction and Symbolic Execution of Deep Neural Networks with Bayesian Approximation of Hidden Features
- A Mapping of Assurance Techniques for Learning Enabled Autonomous Systems to the Systems Engineering Lifecycle
- Understanding the Nature of System-Related Issues in Machine Learning Frameworks: An Exploratory Study
- Accelerating Robustness Verification of Deep Neural Networks Guided by Target Labels
- Adversarial Robustness of Deep Learning: Theory, Algorithms, and Applications
- Path Analysis for Effective Fault Localization in Deep Neural Networks
- A proposal and assessment of an improved heuristic for the Eager Test smell detection
- DeepFault: Fault Localization for Deep Neural Networks
- Automatic Fairness Testing of Neural Classifiers through Adversarial Sampling
- Adversarial Defense Through Network Profiling Based Path Extraction
- NeuralVis: Visualizing and Interpreting Deep Learning Models
- DeepPayload: Black-box Backdoor Attack on Deep Learning Models through Neural Payload Injection
- Automated Testing for Deep Learning Systems with Differential Behavior Criteria
- Automatically Detecting Numerical Instability in Machine Learning Applications via Soft Assertions
- Simulator Ensembles for Trustworthy Autonomous Driving Systems Testing
- Quality Management of Machine Learning Systems
- Explaining Image Classifiers using Statistical Fault Localization
- Structure-Invariant Testing for Machine Translation
- Corner case data description and detection
- Towards Interpreting Recurrent Neural Networks through Probabilistic Abstraction
- Attack as Defense: Characterizing Adversarial Examples using Robustness
- Testing Deep Learning Models: A First Comparative Study of Multiple Testing Techniques
- Fine Grained Dataflow Tracking with Proximal Gradients
- VeriDL: Integrity Verification of Outsourced Deep Learning Services (Extended Version)
- Towards Structured Evaluation of Deep Neural Network Supervisors
- Towards Adversarial Configurations for Software Product Lines
- ROBY: Evaluating the Robustness of a Deep Model by its Decision Boundaries
- In-Simulation Testing of Deep Learning Vision Models in Autonomous Robotic Manipulators
- DeepSmartFuzzer: Reward Guided Test Generation For Deep Learning
- SOCRATES: Towards a Unified Platform for Neural Network Analysis
- Reliability Validation of Learning Enabled Vehicle Tracking
- nn-dependability-kit: Engineering Neural Networks for Safety-Critical Autonomous Driving Systems
- Failure-Scenario Maker for Rule-Based Agent using Multi-agent Adversarial Reinforcement Learning and its Application to Autonomous Driving
- Black-box Adversarial Sample Generation Based on Differential Evolution
- Grand Challenges in Resilience: Autonomous System Resilience through Design and Runtime Measures
- Exposing Semantic Segmentation Failures via Maximum Discrepancy Competition
- An Empirical Study of Bugs in Data Visualization Libraries
- Metamorphic Testing of a Deep Learning based Forecaster
- Generating Adversarial Inputs Using A Black-box Differential Technique
- DeepRepair: Style-Guided Repairing for DNNs in the Real-world Operational Environment
- Formal Verification of Neural Network Controlled Autonomous Systems
- Validate and Enable Machine Learning in Industrial AI
- Towards Quality Assurance of Software Product Lines with Adversarial Configurations
- Boosting the Robustness Verification of DNN by Identifying the Achilles's Heel
- How to Learn a Model Checker
- On Functional Test Generation for Deep Neural Network IPs
- Graph-Based Fuzz Testing for Deep Learning Inference Engine
- Efficient Adversarial Input Generation via Neural Net Patching
- Self-Checking Deep Neural Networks in Deployment
- Active Assessment of Prediction Services as Accuracy Surface Over Attribute Combinations
- Identifying phase transitions in physical systems with neural networks: a neural architecture search perspective
- Debiased Subjective Assessment of Real-World Image Enhancement
- NEUROSPF: A tool for the Symbolic Analysis of Neural Networks
- Testing Machine Translation via Referential Transparency
- SINVAD: Search-based Image Space Navigation for DNN Image Classifier Test Input Generation
- Learning a Safety Verifiable Adaptive Cruise Controller from Human Driving Data
- Mutation Testing framework for Machine Learning
- Testing Autonomous Systems with Believed Equivalence Refinement
- HAWKEYE: Adversarial Example Detector for Deep Neural Networks
- Detecting Deep Neural Network Defects with Data Flow Analysis
- Performance Analysis of Out-of-Distribution Detection on Various Trained Neural Networks
- Security Analysis of Capsule Network Inference using Horizontal Collaboration
- Distortion and Faults in Machine Learning Software
- Neuron Coverage-Guided Domain Generalization
- Logically Sound Arguments for the Effectiveness of ML Safety Measures
- MetaSpace: Metamorphic Testing for Spatial Cognition in Embodied Agents
- On the Mistaken Assumption of Interchangeable Deep Reinforcement Learning Implementations