6 papers
LLM Forensics: Where Do Backdoors Hide? Localizing and Controlling Trigger Mechanisms with Sparse Autoencoders
Wissam Antoun, Francis Kulumba, Théo Lasnier +2
Even though backdoors in LLMs have been a growing concern, their inner workings are still under heavy scrutiny. Trigger-based backdoors are easy to define behaviorally, a rare inpu…
Where Does Authorship Signal Emerge in Encoder-Based Language Models?
Francis Kulumba, Guillaume Vimont, Laurent Romary +1
Authorship attribution models fine-tuned with the same pretrained encoder, data, and loss can differ four-fold in performance depending only on their scoring mechanism. We use mech…
Language-Switching Triggers Take a Latent Detour Through Language Models
Francis Kulumba, Wissam Antoun, Théo Lasnier +2
Backdoor attacks on language models pose a growing security concern, yet the internal mechanisms by which a trigger sequence hijacks model computations remain poorly understood. We…
Language Triggers Hijack Language Circuits: A Mechanistic Analysis of Backdoor Behaviors in Large Language Models
Théo Lasnier, Wissam Antoun, Francis Kulumba +2
Backdoor attacks pose significant security risks for Large Language Models (LLMs), yet the internal mechanisms by which triggers operate remain poorly understood. We present the fi…
CamemBERT 2.0: A Smarter French Language Model Aged to Perfection
Wissam Antoun, Francis Kulumba, Rian Touchent +3
French language models, such as CamemBERT, have been widely adopted across industries for natural language processing (NLP) tasks, with models like CamemBERT seeing over 4 million…
HALvest-Contrastive: Retrieval-Like Authorship Attribution with Patch-Level Late Interaction
Francis Kulumba, Wissam Antoun, Guillaume Vimont +2
Authorship attribution asks whether two pieces of text share a writer, but topical confound makes the task deceptively easy: two authors covering the same topic may look more alike…