4 papers
Backdoor Channels Hidden in Latent Space: Extending Cryptographic Undetectability to Modern Neural Networks
Marte Eggen, Eirik Reiestad, Kristian Gjøsteen +1
Recent cryptographic results establish that neural networks can be backdoored such that no efficient algorithm can distinguish them from a clean model. These guarantees, however, h…
Probing the Probes: Methods and Metrics for Concept Alignment
Jacob Lysnæs-Larsen, Marte Eggen, Inga Strümke
In explainable AI, Concept Activation Vectors (CAVs) are typically obtained by training linear classifier probes to detect human-understandable concepts as directions in the activa…
Integrating attention into explanation frameworks for language and vision transformers
Marte Eggen, Jacob Lysnæs-Larsen, Inga Strümke
The attention mechanism lies at the core of the transformer architecture, providing an interpretable model-internal signal that has motivated a growing interest in attention-based…
A transformer-based deep reinforcement learning approach to spatial navigation in a partially observable Morris Water Maze
Marte Eggen, Inga Strümke
Navigation is a fundamental cognitive skill extensively studied in neuroscientific experiments and has lately gained substantial interest in artificial intelligence research. Recre…