Interpretability Beyond Feature Attribution: Quantitative Testing with Concept Activation Vectors (TCAV)
arXiv:1711.11279
Abstract
The interpretation of deep learning models is a challenge due to their size, complexity, and often opaque internal state. In addition, many systems, such as image classifiers, operate on low-level features rather than high-level concepts. To address these challenges, we introduce Concept Activation Vectors (CAVs), which provide an interpretation of a neural net's internal state in terms of human-friendly concepts. The key idea is to view the high-dimensional internal state of a neural net as an aid, not an obstacle. We show how to use CAVs as part of a technique, Testing with CAVs (TCAV), that uses directional derivatives to quantify the degree to which a user-defined concept is important to a classification result--for example, how sensitive a prediction of "zebra" is to the presence of stripes. Using the domain of image classification as a testing ground, we describe how CAVs may be used to explore hypotheses and generate insights for a standard image classification network as well as a medical application.
References in corpus (8)
- Towards A Rigorous Science of Interpretable Machine Learning
- SmoothGrad: removing noise by adding noise
- Grader variability and the importance of reference standards for evaluating machine learning models for diabetic retinopathy
- Adversarial Machine Learning at Scale
- Grad-CAM: Why did you say that?
- Real Time Image Saliency for Black Box Classifiers
- Interpretation of Neural Networks is Fragile
- ConvNets and ImageNet Beyond Accuracy: Understanding Mistakes and Uncovering Biases
Cited by in corpus (92)
- DeepGauge: Multi-Granularity Testing Criteria for Deep Learning Systems
- Towards Explainable Artificial Intelligence
- Deep k-Nearest Neighbors: Towards Confident, Interpretable and Robust Deep Learning
- Concept Whitening for Interpretable Image Recognition
- Post-hoc Interpretability for Neural NLP: A Survey
- Dermatologist-like explainable AI enhances trust and confidence in diagnosing melanoma
- From Attribution Maps to Human-Understandable Explanations through Concept Relevance Propagation
- Visualizing and Measuring the Geometry of BERT
- Scalable agent alignment via reward modeling: a research direction
- Beyond Expertise and Roles: A Framework to Characterize the Stakeholders of Interpretable Machine Learning and their Needs
- Acquisition of Chess Knowledge in AlphaZero
- Tag N' Train: A Technique to Train Improved Classifiers on Unlabeled Data
- MetaPoison: Practical General-purpose Clean-label Data Poisoning
- How can I choose an explainer? An Application-grounded Evaluation of Post-hoc Explanations
- Explainable Artificial Intelligence Approaches: A Survey
- Using Sequences of Life-events to Predict Human Lives
- Overlearning Reveals Sensitive Attributes
- Explainability in Music Recommender Systems
- Formalizing Trust in Artificial Intelligence: Prerequisites, Causes and Goals of Human Trust in AI
- Neuron Shapley: Discovering the Responsible Neurons
- Towards Relatable Explainable AI with the Perceptual Process
- Facilitating Knowledge Sharing from Domain Experts to Data Scientists for Building NLP Models
- On Interpretability of Artificial Neural Networks: A Survey
- Explaining Explanations: Axiomatic Feature Interactions for Deep Networks
- Survey of XAI in digital pathology
- Evaluating Explanation Without Ground Truth in Interpretable Machine Learning
- Interpretability of a Deep Learning Model in the Application of Cardiac MRI Segmentation with an ACDC Challenge Dataset
- How explainable AI affects human performance: A systematic review of the behavioural consequences of saliency maps
- Considerations When Learning Additive Explanations for Black-Box Models
- Neural network interpretation using descrambler groups
- AI Safety for Everyone
- Evaluating Saliency Map Explanations for Convolutional Neural Networks: A User Study
- Towards falsifiable interpretability research
- Kandinsky Patterns
- Scientific Inference With Interpretable Machine Learning: Analyzing Models to Learn About Real-World Phenomena
- When and How to Fool Explainable Models (and Humans) with Adversarial Examples
- Towards Quantification of Explainability in Explainable Artificial Intelligence Methods
- Improving deep learning with prior knowledge and cognitive models: A survey on enhancing explainability, adversarial robustness and zero-shot learning
- MoËT: Mixture of Expert Trees and its Application to Verifiable Reinforcement Learning
- Unifying machine learning and quantum chemistry -- a deep neural network for molecular wavefunctions
- Explaining Knowledge Distillation by Quantifying the Knowledge
- Domain Knowledge Aided Explainable Artificial Intelligence for Intrusion Detection and Response
- Deeply Explain CNN via Hierarchical Decomposition
- Quantitative Evaluations on Saliency Methods: An Experimental Study
- What Did You Think Would Happen? Explaining Agent Behaviour Through Intended Outcomes
- One Map Does Not Fit All: Evaluating Saliency Map Explanation on Multi-Modal Medical Images
- Evaluating the Stability of Semantic Concept Representations in CNNs for Robust Explainability
- Explaining Explanations to Society
- The AI-DEC: A Card-based Design Method for User-centered AI Explanations
- Expressive Explanations of DNNs by Combining Concept Analysis with ILP
- MiMICRI: Towards Domain-centered Counterfactual Explanations of Cardiovascular Image Classification Models
- Optimising Knee Injury Detection with Spatial Attention and Validating Localisation Ability
- Do Deep Neural Networks Forget Facial Action Units? -- Exploring the Effects of Transfer Learning in Health Related Facial Expression Recognition
- RelatIF: Identifying Explanatory Training Examples via Relative Influence
- SUBPLEX: Towards a Better Understanding of Black Box Model Explanations at the Subpopulation Level
- Concept Tree: High-Level Representation of Variables for More Interpretable Surrogate Decision Trees
- Investigating Bias in Image Classification using Model Explanations
- Towards a Unified Evaluation of Explanation Methods without Ground Truth
- How to Train your Antivirus: RL-based Hardening through the Problem-Space
- RoCourseNet: Distributionally Robust Training of a Prediction Aware Recourse Model
- A Game-Theoretic Taxonomy of Visual Concepts in DNNs
- Verification of Size Invariance in DNN Activations using Concept Embeddings
- The Thousand Faces of Explainable AI Along the Machine Learning Life Cycle: Industrial Reality and Current State of Research
- DeepSeer: Interactive RNN Explanation and Debugging via State Abstraction
- Human-in-the-loop Extraction of Interpretable Concepts in Deep Learning Models
- Towards Robust Metrics for Concept Representation Evaluation
- On the Value of Labeled Data and Symbolic Methods for Hidden Neuron Activation Analysis
- Abstracting Deep Neural Networks into Concept Graphs for Concept Level Interpretability
- Understanding the Dependence of Perception Model Competency on Regions in an Image
- Explainable AI and Adoption of Financial Algorithmic Advisors: an Experimental Study
- Quantifying Learnability and Describability of Visual Concepts Emerging in Representation Learning
- Local Concept Embeddings for Analysis of Concept Distributions in Vision DNN Feature Spaces
- Learning Channel Importance for High Content Imaging with Interpretable Deep Input Channel Mixing
- On convex decision regions in deep network representations
- Explainability Through Systematicity: The Hard Systematicity Challenge for Artificial Intelligence
- A Knowledge Driven Approach to Adaptive Assistance Using Preference Reasoning and Explanation
- DeepVA: Bridging Cognition and Computation through Semantic Interaction and Deep Learning
- Identity Preserve Transform: Understand What Activity Classification Models Have Learnt
- Adversarial TCAV -- Robust and Effective Interpretation of Intermediate Layers in Neural Networks
- Example-Based Concept Analysis Framework for Deep Weather Forecast Models
- Minimal Sufficient Views: A DNN model making predictions with more evidence has higher accuracy
- MEME: Generating RNN Model Explanations via Model Extraction
- Disentangled Neural Architecture Search
- DeepRepViz: Identifying Confounders in Deep Learning Model Predictions
- Extracting Interpretable Concept-Based Decision Trees from CNNs
- On the Role of Domain Experts in Creating Effective Tutoring Systems
- Explaining AI-based Decision Support Systems using Concept Localization Maps
- Finding Concept Representations in Neural Networks with Self-Organizing Maps
- Extractive Adversarial Networks: High-Recall Explanations for Identifying Personal Attacks in Social Media Posts
- Knowledge graphs for empirical concept retrieval
- HOLMES: HOLonym-MEronym based Semantic inspection for Convolutional Image Classifiers
- CACTUS: Detecting and Resolving Conflicts in Objective Functions