3 papers
cs.LG2025
Evaluate with the Inverse: Efficient Approximation of Latent Explanation Quality Distribution
Carlos Eiras-Franco, Anna Hedström, Marina M. -C. Höhne
Obtaining high-quality explanations of a model's output enables developers to identify and correct biases, align the system's behavior with human values, and ensure ethical complia…
cs.LG2024
CoSy: Evaluating Textual Explanations of Neurons
Laura Kopf, Philine Lou Bommer, Anna Hedström +3
A crucial aspect of understanding the complex nature of Deep Neural Networks (DNNs) is the ability to explain learned concepts within their latent representations. While methods ex…
cs.LG2024
Quanda: An Interpretability Toolkit for Training Data Attribution Evaluation and Beyond
Dilyara Bareeva, Galip Ãmit Yolcu, Anna Hedström +4
In recent years, training data attribution (TDA) methods have emerged as a promising direction for the interpretability of neural networks. While research around TDA is thriving, l…