5 papers
Interpretability-Guided Soft Pruning of Attention Heads in Vision Transformers
Kamil KsiÄ Å¼ek, Piotr SuszyÅski, MichaÅ Jan WÅodarczyk +2
Vision foundation models, such as DINOv2, learn highly expressive representations but rely on massive, opaque architectures that demand substantial computational power and memory.…
Your CLIP has 164 dimensions of noise: Exploring the embeddings covariance eigenspectrum of contrastively pretrained vision-language transformers
Jakub Grzywaczewski, Dawid PÅudowski, PrzemysÅaw Biecek
Contrastively pre-trained Vision-Language Models (VLMs) serve as powerful feature extractors. Yet, their shared latent spaces are prone to structural anomalies and act as repositor…
LINE: LLM-based Iterative Neuron Explanations for Vision Models
Vladimir Zaigrajew, MichaÅ Piechota, Gaspar Sekula +2
Interpreting individual neurons in deep neural networks is a crucial step towards understanding their complex decision-making processes and ensuring AI safety. Despite recent progr…
SwordBench: Evaluating Orthogonality of Steering Image Representations
Vladimir Zaigrajew, Dawid Pludowski, Hubert Baniecki +1
Steering or intervening on model representations at inference time to correct predictions is essential for AI interpretability and safety, yet existing evaluation protocols are lim…
Attributions All the Way Down? The Metagame of Interpretability
Hubert Baniecki, Przemyslaw Biecek, Fabian Fumagalli
We introduce the metagame, a conceptual framework for quantifying second-order interaction effects of model explanations. For any first-order attribution explaining a model…