Visual Explanation for Deep Metric Learning
arXiv:1909.12977
Abstract
This work explores the visual explanation for deep metric learning and its applications. As an important problem for learning representation, metric learning has attracted much attention recently, while the interpretation of such model is not as well studied as classification. To this end, we propose an intuitive idea to show where contributes the most to the overall similarity of two input images by decomposing the final activation. Instead of only providing the overall activation map of each image, we propose to generate point-to-point activation intensity between two images so that the relationship between different regions is uncovered. We show that the proposed framework can be directly deployed to a large range of metric learning applications and provides valuable information for understanding the model. Furthermore, our experiments show its effectiveness on two potential applications, i.e. cross-view pattern discovery and interactive retrieval. The source code is available at \url{https://github.com/Jeff-Zilence/Explain_Metric_Learning}.
References in corpus (6)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Grad-CAM++: Improved Visual Explanations for Deep Convolutional Networks
- Striving for Simplicity: The All Convolutional Net
- Learning Face Representation from Scratch
- Sanity Checks for Saliency Maps
- VIGOR: Cross-View Image Geo-localization beyond One-to-one Retrieval
Cited by in corpus (6)
- Survey on the Analysis and Modeling of Visual Kinship: A Decade in the Making
- VIGOR: Cross-View Image Geo-localization beyond One-to-one Retrieval
- Adapting Grad-CAM for Embedding Networks
- Revisiting Street-to-Aerial View Image Geo-localization and Orientation Estimation
- Automatic Face Understanding: Recognizing Families in Photos
- Axiomatic Explanations for Visual Search, Retrieval, and Similarity Learning