How to Certify Machine Learning Based Safety-critical Systems? A Systematic Literature Review
arXiv:2107.12045 · doi:10.1007/s10515-022-00337-x
Abstract
Context: Machine Learning (ML) has been at the heart of many innovations over the past years. However, including it in so-called 'safety-critical' systems such as automotive or aeronautic has proven to be very challenging, since the shift in paradigm that ML brings completely changes traditional certification approaches. Objective: This paper aims to elucidate challenges related to the certification of ML-based safety-critical systems, as well as the solutions that are proposed in the literature to tackle them, answering the question 'How to Certify Machine Learning Based Safety-critical Systems?'. Method: We conduct a Systematic Literature Review (SLR) of research papers published between 2015 to 2020, covering topics related to the certification of ML systems. In total, we identified 217 papers covering topics considered to be the main pillars of ML certification: Robustness, Uncertainty, Explainability, Verification, Safe Reinforcement Learning, and Direct Certification. We analyzed the main trends and problems of each sub-field and provided summaries of the papers extracted. Results: The SLR results highlighted the enthusiasm of the community for this subject, as well as the lack of diversity in terms of datasets and type of models. It also emphasized the need to further develop connections between academia and industries to deepen the domain study. Finally, it also illustrated the necessity to build connections between the above mention main pillars that are for now mainly studied separately. Conclusion: We highlighted current efforts deployed to enable the certification of ML based software systems, and discuss some future research directions.
60 pages (92 pages with references and complements), submitted to a journal (Automated Software Engineering). Changes: Emphasizing difference traditional software engineering / ML approach. Adding Related Works, Threats to Validity and Complementary Materials. Adding a table listing papers reference for each section/subsections
References in corpus (19)
- Explaining and Harnessing Adversarial Examples
- DeepStack: Expert-Level Artificial Intelligence in No-Limit Poker
- Metamorphic Testing: A New Approach for Generating Next Test Cases
- Unsolved Problems in ML Safety
- Verifiably Safe Off-Model Reinforcement Learning
- Systematic Testing of Convolutional Neural Networks for Autonomous Driving
- Worst Cases Policy Gradients
- Can We Trust You? On Calibration of a Probabilistic Object Detector for Autonomous Driving
- Experience-Based Heuristic Search: Robust Motion Planning with Deep Q-Learning
- Coverage Testing of Deep Learning Models using Dataset Characterization
- Under the Hood of Neural Networks: Characterizing Learned Representations by Functional Neuron Populations and Network Ablations
- Certification of embedded systems based on Machine Learning: A survey
- Playing it Safe: Adversarial Robustness with an Abstain Option
- Improved Adversarial Robustness via Logit Regularization Methods
- Unsupervised Data Uncertainty Learning in Visual Retrieval Systems
- GraN: An Efficient Gradient-Norm Based Detector for Adversarial and Misclassified Examples
- Brain-inspired reverse adversarial examples
- How Do You Act? An Empirical Study to Understand Behavior of Deep Reinforcement Learning Agents
- Affine Disentangled GAN for Interpretable and Robust AV Perception
Cited by in corpus (8)
- Hardware Approximate Techniques for Deep Neural Network Accelerators: A Survey
- DeepGD: A Multi-Objective Black-Box Test Selection Approach for Deep Neural Networks
- Optimality Principles in Spacecraft Neural Guidance and Control
- A Probabilistic Framework for Mutation Testing in Deep Neural Networks
- Perception Simplex: Verifiable Collision Avoidance in Autonomous Vehicles Amidst Obstacle Detection Faults
- Synergistic Perception and Control Simplex for Verifiable Safe Vertical Landing
- Bayes2IMC: In-Memory Computing for Bayesian Binary Neural Networks
- Policy Testing with MDPFuzz (Replicability Study)