8 papers
Playing the network backward: A Game Theoretic Attribution Framework
Jakob Paul Zimmermann, Jim Berend, Georg Loho +2
Attribution methods explain which input features drive a model's prediction, making them central to model debugging and mechanistic interpretability. Yet backward attribution metho…
Contrastive Semantic Projection: Faithful Neuron Labeling with Contrastive Examples
Oussama Bouanani, Jim Berend, Wojciech Samek +2
Neuron labeling assigns textual descriptions to internal units of deep networks. Existing approaches typically rely on highly activating examples, often yielding broad or misleadin…
Atlas-Alignment: Making Interpretability Transferable Across Language Models
Bruno Puri, Jim Berend, Sebastian Lapuschkin +1
Interpretability is crucial for building safe, reliable, and controllable language models, yet existing interpretability pipelines remain costly and difficult to scale. Interpretin…
X-SYS: A Reference Architecture for Interactive Explanation Systems
Tobias Labarta, Nhi Hoang, Maximilian Dreyer +5
The explainable AI (XAI) research community has proposed numerous technical methods, yet deploying explainability as systems remains challenging: Interactive explanation systems re…
Efficient and Flexible Neural Network Training through Layer-wise Feedback Propagation
Leander Weber, Jim Berend, Moritz Weckbecker +4
Gradient-based optimization has been a cornerstone of machine learning that enabled the vast advances of Artificial Intelligence (AI) development over the past decades. However, th…
From What to How: Attributing CLIP's Latent Components Reveals Unexpected Semantic Reliance
Maximilian Dreyer, Lorenz Hufe, Jim Berend +3
Transformer-based CLIP models are widely used for text-image probing and feature extraction, making it relevant to understand the internal mechanisms behind their predictions. Whil…