3 citations · 3 across the 5 of their papers we have counts for
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Influence Score and Transformers interpretability: Measure of the Effective Impact of Attention Heads at inference time
Lisa Bouger, Yannick Teglia, Philippe Loubet Moundi
We propose an influence score to quantify the contribution of attention heads to classification decisions in Transformer-based models designed for prompt injection detection. The s…
cs.CL2026
Backdoor Unlearning Generalization: A Path Toward the Removal of Unknown Triggers in LLMs
Lisa Bouger, Théo Lasnier, Philippe Loubet Moundi +2
Backdoor attacks in Large Language Models (LLMs) are a growing security concern, where models can generate adversary-chosen content. Existing defenses target backdoors one at a tim…