4 papers
Conceptualizing Embeddings: Sparse Disentanglement for Vision-Language Models
Piotr Kubaty, Patryk MarszaÅek, Åukasz Struski +3
Vision-language models learn powerful multimodal embeddings, yet their internal semantics remain opaque. While sparse autoencoders (SAEs) can extract interpretable features, they r…
DAVE: Distribution-aware Attribution via ViT Gradient Decomposition
Adam Wróbel, Siddhartha Gairola, Jacek Tabor +3
Vision Transformers (ViTs) have become a dominant architecture in computer vision, yet producing stable and high-resolution attribution maps for these models remains challenging. A…
ProtoQuant: Quantization of Prototypical Parts For General and Fine-Grained Image Classification
MikoÅaj Janusz, Adam Wróbel, Bartosz ZieliÅski +1
Prototypical parts-based models offer a "this looks like that" paradigm for intrinsic interpretability, yet they typically struggle with ImageNet-scale generalization and often req…
Personalized Interpretability -- Interactive Alignment of Prototypical Parts Networks
Tomasz Michalski, Adam Wróbel, Andrea Bontempelli +6
Concept-based interpretable neural networks have gained significant attention due to their intuitive and easy-to-understand explanations based on case-based reasoning, such as "thi…