papers

Publications (14)

cs.LG2025

Counterfactual Influence as a Distributional Quantity

Matthieu Meeus, Igor Shilov, Georgios Kaissis +1

Machine learning models are known to memorize samples from their training data, raising concerns around privacy and generalization. Counterfactual self-influence is a popular metri…

cs.CL2024

ChocoLlama: Lessons Learned From Teaching Llamas Dutch

Matthieu Meeus, Anthony Rathé, François Remy +3

While Large Language Models (LLMs) have shown remarkable capabilities in natural language understanding and generation, their performance often lags in lower-resource, non-English…

cs.CR2023

Achilles' Heels: Vulnerable Record Identification in Synthetic Data Publishing

Matthieu Meeus, Florent Guépin, Ana-Maria Cretu +1

Synthetic data is seen as the most promising solution to share individual-level data while preserving privacy. Shadow modeling-based Membership Inference Attacks (MIAs) have become…

cs.CL2025

The Canary's Echo: Auditing Privacy Risks of LLM-Generated Synthetic Text

Matthieu Meeus, Lukas Wutschitz, Santiago Zanella-Béguelin +2

How much information about training samples can be leaked through synthetic data generated by Large Language Models (LLMs)? Overlooking the subtleties of information flow in synthe…

cs.CR2025

Checkpoint-GCG: Auditing and Attacking Fine-Tuning-Based Prompt Injection Defenses

Xiaoxue Yang, Bozhidar Stevanoski, Matthieu Meeus +1

Large language models (LLMs) are increasingly deployed in real-world applications ranging from chatbots to agentic systems, where they are expected to process untrusted data and fo…

cs.CR2023

Synthetic is all you need: removing the auxiliary data assumption for membership inference attacks against synthetic data

Florent Guépin, Matthieu Meeus, Ana-Maria Cretu +1

Synthetic data is emerging as one of the most promising solutions to share individual-level data while safeguarding privacy. While membership inference attacks (MIAs), based on sha…