12 citations · 28 across the 8 of their papers we have counts for
12 papers
FeatInv: Spatially resolved mapping from feature space to input space using conditional diffusion models
Nils Neukirch, Johanna Vielhaben, Nils Strodthoff
Internal representations are crucial for understanding deep neural networks, such as their properties and reasoning patterns, but remain difficult to interpret. While mapping from…
Mechanistic understanding and validation of large AI models with SemanticLens
Maximilian Dreyer, Jim Berend, Tobias Labarta +4
Unlike human-engineered systems such as aeroplanes, where each component's role and dependencies are well understood, the inner workings of AI models remain largely opaque, hinderi…
Beyond Scalars: Concept-Based Alignment Analysis in Vision Transformers
Johanna Vielhaben, Dilyara Bareeva, Jim Berend +2
Vision transformers (ViTs) can be trained using various learning paradigms, from fully supervised to self-supervised. Diverse training protocols often result in significantly diffe…
PURE: Turning Polysemantic Neurons Into Pure Features by Identifying Relevant Circuits
Maximilian Dreyer, Erblina Purelku, Johanna Vielhaben +2
The field of mechanistic interpretability aims to study the role of individual neurons in Deep Neural Networks. Single neurons, however, have the capability to act polysemantically…
Decoupling Pixel Flipping and Occlusion Strategy for Consistent XAI Benchmarks
Stefan Blücher, Johanna Vielhaben, Nils Strodthoff
Feature removal is a central building block for eXplainable AI (XAI), both for occlusion-based explanations (Shapley values) as well as their evaluation (pixel flipping, PF). Howev…
XAI-based Comparison of Input Representations for Audio Event Classification
Annika Frommholz, Fabian Seipel, Sebastian Lapuschkin +2
Deep neural networks are a promising tool for Audio Event Classification. In contrast to other data like natural images, there are many sensible and non-obvious representations for…