MAGIX: Model Agnostic Globally Interpretable Explanations
arXiv:1706.07160
Abstract
Explaining the behavior of a black box machine learning model at the instance level is useful for building trust. However, it is also important to understand how the model behaves globally. Such an understanding provides insight into both the data on which the model was trained and the patterns that it learned. We present here an approach that learns if-then rules to globally explain the behavior of black box machine learning models that have been used to solve classification problems. The approach works by first extracting conditions that were important at the instance level and then evolving rules through a genetic algorithm with an appropriate fitness function. Collectively, these rules represent the patterns followed by the model for decisioning and are useful for understanding its behavior. We demonstrate the validity and usefulness of the approach by interpreting black box models created using publicly available data sets as well as a private digital marketing data set.
References in corpus (9)
- Towards A Rigorous Science of Interpretable Machine Learning
- Axiomatic Attribution for Deep Networks
- Understanding Black-box Predictions via Influence Functions
- Grad-CAM: Why did you say that?
- Interpreting Blackbox Models via Model Extraction
- An unexpected unity among methods for interpreting model predictions
- Interpreting Classifiers through Attribute Interactions in Datasets
- Interpretable Active Learning
- Human Understandable Explanation Extraction for Black-box Classification Models Based on Matrix Factorization
Cited by in corpus (10)
- Interpretable Machine Learning -- A Brief History, State-of-the-Art and Challenges
- Explanations of model predictions with live and breakDown packages
- Explaining Explanations: Axiomatic Feature Interactions for Deep Networks
- DALEX: explainers for complex predictive models
- Why model why? Assessing the strengths and limitations of LIME
- SAFE ML: Surrogate Assisted Feature Extraction for Model Learning
- Lifting Interpretability-Performance Trade-off via Automated Feature Engineering
- Distilling neural networks into skipgram-level decision lists
- On the Use of Interpretable Machine Learning for the Management of Data Quality
- Interpretability of Blackbox Machine Learning Models through Dataview Extraction and Shadow Model creation