Toward a Unified Framework for Debugging Concept-based Models
arXiv:2109.11160
Abstract
In this paper, we tackle interactive debugging of "gray-box" concept-based models (CBMs). These models learn task-relevant concepts appearing in the inputs and then compute a prediction by aggregating the concept activations. Our work stems from the observation that in CBMs both the concepts and the aggregation function can be affected by different kinds of bugs, and that fixing these bugs requires different kinds of corrective supervision. To this end, we introduce a simple schema for human supervisors to identify and prioritize bugs in both components, and discuss solution strategies and open problems. We also introduce a novel loss function for debugging the aggregation step that generalizes existing strategies for aligning black-box models to CBMs by making them robust to how the concepts change during training.
11 pages, 1 figure. Accepted at the AAAI-22 Workshop on Interactive Machine Learning
References in corpus (11)
- Knowledge Distillation: A Survey
- Concept Whitening for Interpretable Image Recognition
- Deep Learning for Case-Based Reasoning through Prototypes: A Neural Network that Explains Its Predictions
- Concept Bottleneck Models
- This Looks Like That, Because ... Explaining Prototypes for Interpretable Image Recognition
- This Looks Like That... Does it? Shortcomings of Latent Space Prototype Interpretability in Deep Networks
- IAIA-BL: A Case-based Interpretable Deep Learning Model for Classification of Mass Lesions in Digital Mammography
- Debiasing Concept-based Explanations with Causal Analysis
- Machine Guides, Human Supervises: Interactive Learning with Global Explanations
- Interactive Label Cleaning with Example-based Explanations
- Learning Interpretable Concept-Based Models with Human Feedback