2 papers
cs.CV2026
Measuring Monosemanticity in Sparse Autoencoders via Latent Activation Coherence
Katarzyna Filus, Sebastian PokuciÅski
Within Explainable Artificial Intelligence, mechanistic interpretability uses Sparse Autoencoders (SAEs) to extract more interpretable features from neural representations. However…
cs.CV2025
Semantically Guided Adversarial Testing of Vision Models Using Language Models
Katarzyna Filus, Jorge M. Cruz-Duarte
In targeted adversarial attacks on vision models, the selection of the target label is a critical yet often overlooked determinant of attack success. This target label corresponds…