activity
20242026
collaborators

8 papers

cs.LG2026

Attributions All the Way Down? The Metagame of Interpretability

Hubert Baniecki, Przemyslaw Biecek, Fabian Fumagalli

We introduce the metagame, a conceptual framework for quantifying second-order interaction effects of model explanations. For any first-order attribution explaining a model…

cs.LG2025

On the Robustness of Global Feature Effect Explanations

Hubert Baniecki, Giuseppe Casalicchio, Bernd Bischl +1

We study the robustness of global post-hoc explanations for predictive models trained on tabular data. Effects of predictor features in black-box supervised learning are an essenti…

cs.CR2025

Adversarial attacks and defenses in explainable artificial intelligence: A survey

Hubert Baniecki, Przemyslaw Biecek

Explainable artificial intelligence (XAI) methods are portrayed as a remedy for debugging and trusting statistical and deep learning models, as well as interpreting their predictio…

cs.CV2025

Interpreting CLIP with Hierarchical Sparse Autoencoders

Vladimir Zaigrajew, Hubert Baniecki, Przemyslaw Biecek

Sparse autoencoders (SAEs) are useful for detecting and steering interpretable features in neural networks, with particular potential for understanding complex multimodal represent…

cs.LG2025

Efficient and Accurate Explanation Estimation with Distribution Compression

Hubert Baniecki, Giuseppe Casalicchio, Bernd Bischl +1

We discover a theoretical connection between explanation estimation and distribution compression that significantly improves the approximation of feature attributions, importance,…

cs.CV2024

Aggregated Attributions for Explanatory Analysis of 3D Segmentation Models

Maciej Chrabaszcz, Hubert Baniecki, Piotr Komorowski +2

Analysis of 3D segmentation models, especially in the context of medical imaging, is often limited to segmentation performance metrics that overlook the crucial aspect of explainab…