Technologies for Trustworthy Machine Learning: A Survey in a Socio-Technical Context
arXiv:2007.08911
Abstract
Concerns about the societal impact of AI-based services and systems has encouraged governments and other organisations around the world to propose AI policy frameworks to address fairness, accountability, transparency and related topics. To achieve the objectives of these frameworks, the data and software engineers who build machine-learning systems require knowledge about a variety of relevant supporting tools and techniques. In this paper we provide an overview of technologies that support building trustworthy machine learning systems, i.e., systems whose properties justify that people place trust in them. We argue that four categories of system properties are instrumental in achieving the policy objectives, namely fairness, explainability, auditability and safety & security (FEAS). We discuss how these properties need to be considered across all stages of the machine learning life cycle, from data collection through run-time model inference. As a consequence, we survey in this paper the main technologies with respect to all four of the FEAS properties, for data-centric as well as model-centric stages of the machine learning system life cycle. We conclude with an identification of open research problems, with a particular focus on the connection between trustworthy machine learning technologies and their implications for individuals and society.
We are updating some sections to include more recent advances
References in corpus (21)
- Towards A Rigorous Science of Interpretable Machine Learning
- Methods for Interpreting and Understanding Deep Neural Networks
- Delving into Transferable Adversarial Examples and Black-box Attacks
- Poisoning Attacks against Support Vector Machines
- Towards Deep Neural Network Architectures Robust to Adversarial Examples
- Security Evaluation of Pattern Classifiers under Attack
- Generating Adversarial Malware Examples for Black-Box Attacks Based on GAN
- InterpretML: A Unified Framework for Machine Learning Interpretability
- Data Decisions and Theoretical Implications when Adversarially Learning Fair Representations
- On the (im)possibility of fairness
- Adversarial Feature Selection against Evasion Attacks
- Defensive Distillation is Not Robust to Adversarial Examples
- Fast Feature Fool: A data independent approach to universal adversarial perturbations
- On the relation between accuracy and fairness in binary classification
- Randomized Prediction Games for Adversarial Machine Learning
- Quantifying Interpretability and Trust in Machine Learning Systems
- A statistical framework for fair predictive algorithms
- Differentially Private Bayesian Optimization
- Orchestrating the Development Lifecycle of Machine Learning-Based IoT Applications: A Taxonomy and Survey
- An Analytical Survey of Provenance Sanitization
- ProvAbs: model, policy, and tooling for abstracting PROV graphs