9 papers
Interpretability Without Tradeoffs: Disentangling Polysemanticity At Equal Predictive Performance
DoÄukan BaÄcı, Bernt Schiele, Simone Schaub-Meyer +2
Deep neural networks (DNNs) are widely used, but interpreting what they actually learn remains difficult. A major obstacle is that individual neurons often encode multiple unrelate…
Certified Circuits: Stability Guarantees for Mechanistic Circuits
Alaa Anani, Tobias Lorenz, Bernt Schiele +2
Understanding how neural networks arrive at their predictions is essential for debugging, auditing, and deployment. Mechanistic interpretability pursues this goal by identifying ci…
Seeing Through Circuits: Faithful Mechanistic Interpretability for Vision Transformers
Nina Żukowska, Wolfgang Stammer, Bernt Schiele +1
Transparency of neural networks' internal reasoning is at the heart of interpretability research, adding to trust, safety, and understanding of these models. The field of mechanist…
FaCT: Faithful Concept Traces for Explaining Neural Network Decisions
Amin Parchami-Araghi, Sukrut Rao, Jonas Fischer +1
Deep networks have shown remarkable performance across a wide range of tasks, yet getting a global concept-level understanding of how they function remains a key challenge. Many po…
CFM: Language-aligned Concept Foundation Model for Vision
Kai Wittenmayer, Sukrut Rao, Amin Parchami-Araghi +2
Language-aligned vision foundation models perform strongly across diverse downstream tasks. Yet, their learned representations remain opaque, making interpreting their decision-mak…
VITAL: More Understandable Feature Visualization through Distribution Alignment and Relevant Information Flow
Ada Gorgun, Bernt Schiele, Jonas Fischer
Neural networks are widely adopted to solve complex and challenging tasks. Especially in high-stakes decision-making, understanding their reasoning process is crucial, yet proves c…