5 papers
Proxy-Based Approximation of Shapley and Banzhaf Interactions
Santo M. A. R. Thies, Hubert Baniecki, R. Teal Witter +3
Shapley and Banzhaf interactions capture the complex dynamics inherent in modern machine learning applications. However, current estimators for these higher-order interactions trad…
SwordBench: Evaluating Orthogonality of Steering Image Representations
Vladimir Zaigrajew, Dawid Pludowski, Hubert Baniecki +1
Steering or intervening on model representations at inference time to correct predictions is essential for AI interpretability and safety, yet existing evaluation protocols are lim…
Functional Decomposition and Shapley Interactions for Interpreting Survival Models
Sophie Hanna Langbein, Hubert Baniecki, Fabian Fumagalli +3
Hazard and survival functions are natural, interpretable targets in time-to-event prediction, but their inherent non-additivity fundamentally limits standard additive explanation m…
Explaining Similarity in Vision-Language Encoders with Weighted Banzhaf Interactions
Hubert Baniecki, Maximilian Muschalik, Fabian Fumagalli +3
Language-image pre-training (LIP) enables the development of vision-language models capable of zero-shot classification, localization, multimodal retrieval, and semantic understand…
Birds look like cars: Adversarial analysis of intrinsically interpretable deep learning
Hubert Baniecki, Przemyslaw Biecek
A common belief is that intrinsically interpretable deep learning models ensure a correct, intuitive understanding of their behavior and offer greater robustness against accidental…