3 papers
cs.CR2026
Backdoor Channels Hidden in Latent Space: Cryptographic Undetectability in Modern Neural Networks
Marte Eggen, Eirik Reiestad, Kristian Gjøsteen +1
Recent cryptographic results establish that neural networks can be backdoored such that no efficient algorithm can distinguish them from a clean model. These guarantees, however, h…
cs.AI2025
Probing the Probes: Methods and Metrics for Concept Alignment
Jacob Lysnæs-Larsen, Marte Eggen, Inga Strümke
In explainable AI, Concept Activation Vectors (CAVs) are typically obtained by training linear classifier probes to detect human-understandable concepts as directions in the activa…
cs.LG2025
Integrating attention into explanation frameworks for language and vision transformers
Marte Eggen, Jacob Lysnæs-Larsen, Inga Strümke
The attention mechanism lies at the core of the transformer architecture, providing an interpretable model-internal signal that has motivated a growing interest in attention-based…