"Why Should You Trust My Explanation?" Understanding Uncertainty in LIME Explanations
arXiv:1904.12991
Abstract
Methods for interpreting machine learning black-box models increase the outcomes' transparency and in turn generates insight into the reliability and fairness of the algorithms. However, the interpretations themselves could contain significant uncertainty that undermines the trust in the outcomes and raises concern about the model's reliability. Focusing on the method "Local Interpretable Model-agnostic Explanations" (LIME), we demonstrate the presence of two sources of uncertainty, namely the randomness in its sampling procedure and the variation of interpretation quality across different input data points. Such uncertainty is present even in models with high training and test accuracy. We apply LIME to synthetic data and two public data sets, text classification in 20 Newsgroup and recidivism risk-scoring in COMPAS, to support our argument.
References in corpus (2)
Cited by in corpus (7)
- Explainability Fact Sheets: A Framework for Systematic Assessment of Explainable Approaches
- To trust or not to trust an explanation: using LEAF to evaluate local linear XAI methods
- Fairness via Explanation Quality: Evaluating Disparities in the Quality of Post hoc Explanations
- A study of data and label shift in the LIME framework
- bLIMEy: Surrogate Prediction Explanations Beyond LIME
- What will it take to generate fairness-preserving explanations?
- Show or Suppress? Managing Input Uncertainty in Machine Learning Model Explanations