8 papers
Attributions All the Way Down? The Metagame of Interpretability
Hubert Baniecki, Przemyslaw Biecek, Fabian Fumagalli
We introduce the metagame, a conceptual framework for quantifying second-order interaction effects of model explanations. For any first-order attribution explaining a model…
On the Robustness of Global Feature Effect Explanations
Hubert Baniecki, Giuseppe Casalicchio, Bernd Bischl +1
We study the robustness of global post-hoc explanations for predictive models trained on tabular data. Effects of predictor features in black-box supervised learning are an essenti…
Adversarial attacks and defenses in explainable artificial intelligence: A survey
Hubert Baniecki, Przemyslaw Biecek
Explainable artificial intelligence (XAI) methods are portrayed as a remedy for debugging and trusting statistical and deep learning models, as well as interpreting their predictio…
Interpreting CLIP with Hierarchical Sparse Autoencoders
Vladimir Zaigrajew, Hubert Baniecki, Przemyslaw Biecek
Sparse autoencoders (SAEs) are useful for detecting and steering interpretable features in neural networks, with particular potential for understanding complex multimodal represent…
Efficient and Accurate Explanation Estimation with Distribution Compression
Hubert Baniecki, Giuseppe Casalicchio, Bernd Bischl +1
We discover a theoretical connection between explanation estimation and distribution compression that significantly improves the approximation of feature attributions, importance,…
Aggregated Attributions for Explanatory Analysis of 3D Segmentation Models
Maciej Chrabaszcz, Hubert Baniecki, Piotr Komorowski +2
Analysis of 3D segmentation models, especially in the context of medical imaging, is often limited to segmentation performance metrics that overlook the crucial aspect of explainab…