2 citations · 2 across the 4 of their papers we have counts for
4 papers
LINE: LLM-based Iterative Neuron Explanations for Vision Models
Vladimir Zaigrajew, Michał Piechota, Gaspar Sekula +2
Interpreting individual neurons in deep neural networks is a crucial step towards understanding their complex decision-making processes and ensuring AI safety. Despite recent progr…
SwordBench: Evaluating Orthogonality of Steering Image Representations
Vladimir Zaigrajew, Dawid Pludowski, Hubert Baniecki +1
Steering or intervening on model representations at inference time to correct predictions is essential for AI interpretability and safety, yet existing evaluation protocols are lim…
Interpreting CLIP with Hierarchical Sparse Autoencoders
Vladimir Zaigrajew, Hubert Baniecki, Przemyslaw Biecek
Sparse autoencoders (SAEs) are useful for detecting and steering interpretable features in neural networks, with particular potential for understanding complex multimodal represent…
Red Teaming Models for Hyperspectral Image Analysis Using Explainable AI
Vladimir Zaigrajew, Hubert Baniecki, Lukasz Tulczyjew +4
Remote sensing (RS) applications in the space domain demand machine learning (ML) models that are reliable, robust, and quality-assured, making red teaming a vital approach for ide…