Testing Deep Neural Networks
arXiv:1803.04792
Abstract
Deep neural networks (DNNs) have a wide range of applications, and software employing them must be thoroughly tested, especially in safety-critical domains. However, traditional software test coverage metrics cannot be applied directly to DNNs. In this paper, inspired by the MC/DC coverage criterion, we propose a family of four novel test criteria that are tailored to structural features of DNNs and their semantics. We validate the criteria by demonstrating that the generated test inputs guided via our proposed coverage criteria are able to capture undesired behaviours in a DNN. Test cases are generated using a symbolic approach and a gradient-based heuristic search. By comparing them with existing methods, we show that our criteria achieve a balance between their ability to find bugs (proxied using adversarial examples) and the computational cost of test case generation. Our experiments are conducted on state-of-the-art DNNs obtained using popular open source datasets, including MNIST, CIFAR-10 and ImageNet.
References in corpus (20)
- Understanding Neural Networks Through Deep Visualization
- Provable defenses against adversarial examples via the convex outer adversarial polytope
- Guiding Deep Learning System Testing using Surprise Adequacy
- Formal Security Analysis of Neural Networks using Symbolic Intervals
- An approach to reachability analysis for feed-forward ReLU neural networks
- Automated Directed Fairness Testing
- Identifying Implementation Bugs in Machine Learning based Image Classifiers using Metamorphic Testing
- TensorFuzz: Debugging Neural Networks with Coverage-Guided Fuzzing
- A Game-Based Approximate Verification of Deep Neural Networks with Provable Guarantees
- Universal adversarial perturbations
- Output Range Analysis for Deep Neural Networks
- DeepRoad: GAN-based Metamorphic Autonomous Driving System Testing
- Combinatorial Testing for Deep Learning Systems
- A Survey of Safety and Trustworthiness of Deep Neural Networks: Verification, Testing, Adversarial Attack and Defence, and Interpretability
- Symbolic Execution for Deep Neural Networks
- Detecting Adversarial Samples for Deep Neural Networks through Mutation Testing
- Systematic Testing of Convolutional Neural Networks for Autonomous Driving
- DeepHunter: Hunting Deep Neural Network Defects via Coverage-Guided Fuzzing
- Global Robustness Evaluation of Deep Neural Networks with Provable Guarantees for the Norm
- A Formalization of Robustness for Deep Neural Networks
Cited by in corpus (46)
- Machine Learning Testing: Survey, Landscapes and Horizons
- Verification for Machine Learning, Autonomy, and Neural Networks Survey
- Combinatorial Testing for Deep Learning Systems
- A Survey of Safety and Trustworthiness of Deep Neural Networks: Verification, Testing, Adversarial Attack and Defence, and Interpretability
- Using Machine Learning Safely in Automotive Software: An Assessment and Adaption of Software Process Requirements in ISO 26262
- DeepGD: A Multi-Objective Black-Box Test Selection Approach for Deep Neural Networks
- DeepCruiser: Automated Guided Testing for Stateful Deep Learning Systems
- Software Testing for Machine Learning
- Increasing the Confidence of Deep Neural Networks by Coverage Analysis
- The RFML Ecosystem: A Look at the Unique Challenges of Applying Deep Learning to Radio Frequency Applications
- Automated Test Generation to Detect Individual Discrimination in AI Models
- Software Engineering Practice in the Development of Deep Learning Applications
- Global Robustness Evaluation of Deep Neural Networks with Provable Guarantees for the Norm
- There is Limited Correlation between Coverage and Robustness for Deep Neural Networks
- Coverage Testing of Deep Learning Models using Dataset Characterization
- RAID: Randomized Adversarial-Input Detection for Neural Networks
- DeepGini: Prioritizing Massive Tests to Enhance the Robustness of Deep Neural Networks
- Correctness Verification of Neural Networks
- Fairness Testing of Deep Image Classification with Adequacy Metrics
- Deep Ensembles from a Bayesian Perspective
- Beyond Accuracy: An Empirical Study on Unit Testing in Open-source Deep Learning Projects
- CAGFuzz: Coverage-Guided Adversarial Generative Fuzzing Testing of Deep Learning Systems
- testRNN: Coverage-guided Testing on Recurrent Neural Networks
- Data Sanity Check for Deep Learning Systems via Learnt Assertions
- Towards Characterizing Adversarial Defects of Deep Learning Software from the Lens of Uncertainty
- Challenges of Testing an Evolving Cancer Registration Support System in Practice
- Revisiting Deep Neural Network Test Coverage from the Test Effectiveness Perspective
- Fixed-Point Code Synthesis For Neural Networks
- DeepSearch: A Simple and Effective Blackbox Attack for Deep Neural Networks
- Path Analysis for Effective Fault Localization in Deep Neural Networks
- DeepFault: Fault Localization for Deep Neural Networks
- Abstraction and Symbolic Execution of Deep Neural Networks with Bayesian Approximation of Hidden Features
- Corner case data description and detection
- Quality Management of Machine Learning Systems
- DeepSaucer: Unified Environment for Verifying Deep Neural Networks
- nn-dependability-kit: Engineering Neural Networks for Safety-Critical Autonomous Driving Systems
- DeepSmartFuzzer: Reward Guided Test Generation For Deep Learning
- Quantitative Projection Coverage for Testing ML-enabled Autonomous Systems
- Explaining Image Classifiers using Statistical Fault Localization
- Attack as Defense: Characterizing Adversarial Examples using Robustness
- Black-box Adversarial Sample Generation Based on Differential Evolution
- Self-Checking Deep Neural Networks in Deployment
- Graph-Based Fuzz Testing for Deep Learning Inference Engine
- On Functional Test Generation for Deep Neural Network IPs
- Boosting the Robustness Verification of DNN by Identifying the Achilles's Heel
- Detecting Deep Neural Network Defects with Data Flow Analysis