7 papers
Interpretability Without Tradeoffs: Disentangling Polysemanticity At Equal Predictive Performance
DoÄukan BaÄcı, Bernt Schiele, Simone Schaub-Meyer +2
Deep neural networks (DNNs) are widely used, but interpreting what they actually learn remains difficult. A major obstacle is that individual neurons often encode multiple unrelate…
What is Missing? Explaining Neurons Activated by Absent Concepts
Robin Hesse, Simone Schaub-Meyer, Janina Hesse +2
Explainable artificial intelligence (XAI) aims to provide human-interpretable insights into the behavior of deep neural networks (DNNs), typically by estimating a simplified causal…
Beyond Accuracy: What Matters in Designing Well-Behaved Image Classification Models?
Robin Hesse, DoÄukan BaÄcı, Bernt Schiele +2
Deep learning has become an essential part of computer vision, with deep neural networks (DNNs) excelling in predictive performance. However, they often fall short in other critica…
Activation Subspaces for Out-of-Distribution Detection
BarıŠZöngür, Robin Hesse, Stefan Roth
To ensure the reliability of deep models in real-world applications, out-of-distribution (OOD) detection methods aim to distinguish samples close to the training distribution (in-d…
Disentangling Polysemantic Channels in Convolutional Neural Networks
Robin Hesse, Jonas Fischer, Simone Schaub-Meyer +1
Mechanistic interpretability is concerned with analyzing individual components in a (convolutional) neural network (CNN) and how they form larger circuits representing decision mec…
Continual Learning Should Move Beyond Incremental Classification
Rupert Mitchell, Antonio Alliegro, Raffaello Camoriano +17
Continual learning (CL) is the sub-field of machine learning concerned with accumulating knowledge in dynamic environments. So far, CL research has mainly focused on incremental cl…