Finding and Fixing Spurious Patterns with Explanations
arXiv:2106.02112
Abstract
Image classifiers often use spurious patterns, such as "relying on the presence of a person to detect a tennis racket, which do not generalize. In this work, we present an end-to-end pipeline for identifying and mitigating spurious patterns for such models, under the assumption that we have access to pixel-wise object-annotations. We start by identifying patterns such as "the model's prediction for tennis racket changes 63% of the time if we hide the people." Then, if a pattern is spurious, we mitigate it via a novel form of data augmentation. We demonstrate that our method identifies a diverse set of spurious patterns and that it mitigates them by producing a model that is both more accurate on a distribution where the spurious pattern is not helpful and more robust to distribution shift.
References in corpus (11)
- Adam: A Method for Stochastic Optimization
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- Deep Inside Convolutional Networks: Visualising Image Classification Models and Saliency Maps
- EdgeConnect: Generative Image Inpainting with Adversarial Edge Learning
- Sanity Checks for Saliency Maps
- Concept Bottleneck Models
- Hierarchical interpretations for neural network predictions
- Model Patching: Closing the Subgroup Performance Gap with Data Augmentation
- Counterfactual Generative Networks
- Causality-aware counterfactual confounding adjustment for feature representations learned by deep models
- Interventional Domain Adaptation