An Objective Metric for Explainable AI: How and Why to Estimate the Degree of Explainability
arXiv:2109.05327 · doi:10.1016/j.knosys.2023.110866
Abstract
Explainable AI was born as a pathway to allow humans to explore and understand the inner working of complex systems. However, establishing what is an explanation and objectively evaluating explainability are not trivial tasks. This paper presents a new model-agnostic metric to measure the Degree of Explainability of information in an objective way. We exploit a specific theoretical model from Ordinary Language Philosophy called the Achinstein's Theory of Explanations, implemented with an algorithm relying on deep language models for knowledge graph extraction and information retrieval. To understand whether this metric can measure explainability, we devised a few experiments and user studies involving more than 190 participants, evaluating two realistic systems for healthcare and finance using famous AI technology, including Artificial Neural Networks and TreeSHAP. The results we obtained are statistically significant (with P values lower than .01), suggesting that our proposed metric for measuring the Degree of Explainability is robust in several scenarios, and it aligns with concrete expectations.
24 pages, 7 figures, 6 tables, Source code available at: https://github.com/Francesco-Sovrano/DoXpy
References in corpus (9)
- A Comprehensive Taxonomy for Explainable Artificial Intelligence: A Systematic Survey of Surveys on Methods and Concepts
- Proxy Tasks and Subjective Measures Can Be Misleading in Evaluating Explainable AI Systems
- Explanation in Human-AI Systems: A Literature Meta-Review, Synopsis of Key Ideas and Publications, and Bibliography for Explainable AI
- Metrics for Explainable AI: Challenges and Prospects
- Explainable AI: Beware of Inmates Running the Asylum Or: How I Learnt to Stop Worrying and Love the Social and Behavioural Sciences
- Generalized Linear Rule Models
- LAReQA: Language-agnostic answer retrieval from a multilingual pool
- Agree to Disagree: Subjective Fairness in Privacy-Restricted Decentralised Conflict Resolution
- QADiscourse -- Discourse Relations as QA Pairs: Representation, Crowdsourcing and Baselines
Cited by in corpus (4)
- An Empirical Study on Compliance with Ranking Transparency in the Software Documentation of EU Online Platforms
- Heterogeneous Subgraph Network with Prompt Learning for Interpretable Depression Detection on Social Media
- XAI-CF -- Examining the Role of Explainable Artificial Intelligence in Cyber Forensics
- Revolutionizing Validation and Verification: Explainable Testing Methodologies for Intelligent Automotive Decision-Making Systems