3 citations · 7 across the 10 of their papers we have counts for
9 papers · 1 filter
ResLRP: The Role of Residual Cancellation in Attribution Instability in Vision Transformers
Jim Berend, Reduan Achtibat, Daniel Schäffer +4
Vision Transformers (ViTs) are central to most modern vision models, yet obtaining input attributions that are fine-grained, faithful, and stable remains challenging. Layer-wise Re…
Concept-based explanation of gene expression prediction from H&E images
Amos Muench, Jonathan Thielmann, Reduan Achtibat +9
Recent advances in pathology foundation models have enabled accurate prediction of spatial transcriptomics (ST) from routine H&E images. However, existing explainability methods fo…
Contrastive Semantic Projection: Faithful Neuron Labeling with Contrastive Examples
Oussama Bouanani, Jim Berend, Wojciech Samek +2
Neuron labeling assigns textual descriptions to internal units of deep networks. Existing approaches typically rely on highly activating examples, often yielding broad or misleadin…
Dyslexify: A Mechanistic Defense Against Typographic Attacks in CLIP
Lorenz Hufe, Constantin Venhoff, Erblina Purelku +3
Typographic attacks exploit multi-modal systems by injecting text into images, leading to targeted misclassifications, malicious content generation and even Vision-Language Model j…
Explainable concept mappings of MRI: Revealing the mechanisms underlying deep learning-based brain disease classification
Christian Tinauer, Anna Damulina, Maximilian Sackl +9
Motivation. While recent studies show high accuracy in the classification of Alzheimer's disease using deep neural networks, the underlying learned concepts have not been investiga…
PURE: Turning Polysemantic Neurons Into Pure Features by Identifying Relevant Circuits
Maximilian Dreyer, Erblina Purelku, Johanna Vielhaben +2
The field of mechanistic interpretability aims to study the role of individual neurons in Deep Neural Networks. Single neurons, however, have the capability to act polysemantically…