25 citations · 33 across the 8 of their papers we have counts for
3 papers · 1 filter
Model Guidance via Explanations Turns Image Classifiers into Segmentation Models
Xiaoyan Yu, Jannik Franzen, Wojciech Samek +2
Heatmaps generated on inputs of image classification networks via explainable AI methods like Grad-CAM and LRP have been observed to resemble segmentations of input images in many…
PURE: Turning Polysemantic Neurons Into Pure Features by Identifying Relevant Circuits
Maximilian Dreyer, Erblina Purelku, Johanna Vielhaben +2
The field of mechanistic interpretability aims to study the role of individual neurons in Deep Neural Networks. Single neurons, however, have the capability to act polysemantically…
Reveal to Revise: An Explainable AI Life Cycle for Iterative Bias Correction of Deep Models
Frederik Pahde, Maximilian Dreyer, Wojciech Samek +1
State-of-the-art machine learning models often learn spurious correlations embedded in the training data. This poses risks when deploying these models for high-stake decision-makin…