3 papers
cs.CV2026
LINE: LLM-based Iterative Neuron Explanations for Vision Models
Vladimir Zaigrajew, MichaÅ Piechota, Gaspar Sekula +2
Interpreting individual neurons in deep neural networks is a crucial step towards understanding their complex decision-making processes and ensuring AI safety. Despite recent progr…
cs.CV2026
SwordBench: Evaluating Orthogonality of Steering Image Representations
Vladimir Zaigrajew, Dawid Pludowski, Hubert Baniecki +1
Steering or intervening on model representations at inference time to correct predictions is essential for AI interpretability and safety, yet existing evaluation protocols are lim…
cs.CV2025
Interpreting CLIP with Hierarchical Sparse Autoencoders
Vladimir Zaigrajew, Hubert Baniecki, Przemyslaw Biecek
Sparse autoencoders (SAEs) are useful for detecting and steering interpretable features in neural networks, with particular potential for understanding complex multimodal represent…