5 papers
When Tables Go Crazy: Evaluating Multimodal Models on French Financial Documents
Virginie Mouilleron, Théo Lasnier, Anna Mosolova +1
Vision-language models (VLMs) perform well on many document understanding tasks, yet their reliability in specialized, non-English domains remains underexplored. This gap is especi…
Backdoor Unlearning Generalization: A Path Toward the Removal of Unknown Triggers in LLMs
Lisa Bouger, Théo Lasnier, Philippe Loubet Moundi +2
Backdoor attacks in Large Language Models (LLMs) are a growing security concern, where models can generate adversary-chosen content. Existing defenses target backdoors one at a tim…
Translation Heads: Disentangling meaning from language in LLM-based machine translation
Théo Lasnier, Armel Zebaze, Djamé Seddah +2
Mechanistic Interpretability (MI) seeks to explain how neural networks implement their capabilities, but the scale of Large Language Models (LLMs) has limited prior MI work in Mach…
Language-Switching Triggers Take a Latent Detour Through Language Models
Francis Kulumba, Wissam Antoun, Théo Lasnier +2
Backdoor attacks on language models pose a growing security concern, yet the internal mechanisms by which a trigger sequence hijacks model computations remain poorly understood. We…
Language Triggers Hijack Language Circuits: A Mechanistic Analysis of Backdoor Behaviors in Large Language Models
Théo Lasnier, Wissam Antoun, Francis Kulumba +2
Backdoor attacks pose significant security risks for Large Language Models (LLMs), yet the internal mechanisms by which triggers operate remain poorly understood. We present the fi…