5 papers
Birds look like cars: Adversarial analysis of intrinsically interpretable deep learning
Hubert Baniecki, Przemyslaw Biecek
A common belief is that intrinsically interpretable deep learning models ensure a correct, intuitive understanding of their behavior and offer greater robustness against accidental…
Adversarial attacks and defenses in explainable artificial intelligence: A survey
Hubert Baniecki, Przemyslaw Biecek
Explainable artificial intelligence (XAI) methods are portrayed as a remedy for debugging and trusting statistical and deep learning models, as well as interpreting their predictio…
Interpreting CLIP with Hierarchical Sparse Autoencoders
Vladimir Zaigrajew, Hubert Baniecki, Przemyslaw Biecek
Sparse autoencoders (SAEs) are useful for detecting and steering interpretable features in neural networks, with particular potential for understanding complex multimodal represent…
Global Counterfactual Directions
Bartlomiej Sobieski, PrzemysÅaw Biecek
Despite increasing progress in development of methods for generating visual counterfactual explanations, especially with the recent rise of Denoising Diffusion Probabilistic Models…
Interpretable Machine Learning for Survival Analysis
Sophie Hanna Langbein, Mateusz KrzyziÅski, MikoÅaj Spytek +3
With the spread and rapid advancement of black box machine learning models, the field of interpretable machine learning (IML) or explainable artificial intelligence (XAI) has becom…