collaborators

5 papers

cs.CV2026

Interpretability-Guided Soft Pruning of Attention Heads in Vision Transformers

Kamil KsiÄ Å¼ek, Piotr Suszyński, Michał Jan Włodarczyk +2

Vision foundation models, such as DINOv2, learn highly expressive representations but rely on massive, opaque architectures that demand substantial computational power and memory.…

cs.CV2026

Your CLIP has 164 dimensions of noise: Exploring the embeddings covariance eigenspectrum of contrastively pretrained vision-language transformers

Jakub Grzywaczewski, Dawid Płudowski, Przemysław Biecek

Contrastively pre-trained Vision-Language Models (VLMs) serve as powerful feature extractors. Yet, their shared latent spaces are prone to structural anomalies and act as repositor…

cs.CV2026

LINE: LLM-based Iterative Neuron Explanations for Vision Models

Vladimir Zaigrajew, Michał Piechota, Gaspar Sekula +2

Interpreting individual neurons in deep neural networks is a crucial step towards understanding their complex decision-making processes and ensuring AI safety. Despite recent progr…

cs.CV2026

SwordBench: Evaluating Orthogonality of Steering Image Representations

Vladimir Zaigrajew, Dawid Pludowski, Hubert Baniecki +1

Steering or intervening on model representations at inference time to correct predictions is essential for AI interpretability and safety, yet existing evaluation protocols are lim…

cs.LG2026

Attributions All the Way Down? The Metagame of Interpretability

Hubert Baniecki, Przemyslaw Biecek, Fabian Fumagalli

We introduce the metagame, a conceptual framework for quantifying second-order interaction effects of model explanations. For any first-order attribution explaining a model…