xGEMs: Generating Examplars to Explain Black-Box Models
arXiv:1806.08867
Abstract
This work proposes xGEMs or manifold guided exemplars, a framework to understand black-box classifier behavior by exploring the landscape of the underlying data manifold as data points cross decision boundaries. To do so, we train an unsupervised implicit generative model -- treated as a proxy to the data manifold. We summarize black-box model behavior quantitatively by perturbing data samples along the manifold. We demonstrate xGEMs' ability to detect and quantify bias in model learning and also for understanding the changes in model behavior as training progresses.
References in corpus (6)
- Towards A Rigorous Science of Interpretable Machine Learning
- Equality of Opportunity in Supervised Learning
- On Calibration of Modern Neural Networks
- SmoothGrad: removing noise by adding noise
- ConvNets and ImageNet Beyond Accuracy: Understanding Mistakes and Uncovering Biases
- Supervised topic models for clinical interpretability
Cited by in corpus (11)
- Getting a CLUE: A Method for Explaining Uncertainty Estimates
- Model-Based Counterfactual Synthesizer for Interpretation
- Evaluating Explainable AI: Which Algorithmic Explanations Help Users Predict Model Behavior?
- Gifsplanation via Latent Shift: A Simple Autoencoder Approach to Counterfactual Generation for Chest X-rays
- Explanation by Progressive Exaggeration
- Dynamic Measurement Scheduling for Event Forecasting using Deep RL
- Selective Classification Can Magnify Disparities Across Groups
- Roadmap of Designing Cognitive Metrics for Explainable Artificial Intelligence (XAI)
- Generative Counterfactuals for Neural Networks via Attribute-Informed Perturbation
- FastIF: Scalable Influence Functions for Efficient Model Interpretation and Debugging
- δ-CLUE: Diverse Sets of Explanations for Uncertainty Estimates