4 papers
Backdoor Channels Hidden in Latent Space: Cryptographic Undetectability in Modern Neural Networks
Marte Eggen, Eirik Reiestad, Kristian Gjøsteen +1
Recent cryptographic results establish that neural networks can be backdoored such that no efficient algorithm can distinguish them from a clean model. These guarantees, however, h…
A Framework for Causal Concept-based Model Explanations
Anna Rodum Bjøru, Jacob Lysnæs-Larsen, Oskar Jørgensen +2
This work presents a conceptual framework for causal concept-based post-hoc Explainable Artificial Intelligence (XAI), based on the requirements that explanations for non-interpret…
Probing the Probes: Methods and Metrics for Concept Alignment
Jacob Lysnæs-Larsen, Marte Eggen, Inga Strümke
In explainable AI, Concept Activation Vectors (CAVs) are typically obtained by training linear classifier probes to detect human-understandable concepts as directions in the activa…
Integrating attention into explanation frameworks for language and vision transformers
Marte Eggen, Jacob Lysnæs-Larsen, Inga Strümke
The attention mechanism lies at the core of the transformer architecture, providing an interpretable model-internal signal that has motivated a growing interest in attention-based…