Publications (14)
Counterfactual Influence as a Distributional Quantity
Matthieu Meeus, Igor Shilov, Georgios Kaissis +1
Machine learning models are known to memorize samples from their training data, raising concerns around privacy and generalization. Counterfactual self-influence is a popular metri…
ChocoLlama: Lessons Learned From Teaching Llamas Dutch
Matthieu Meeus, Anthony Rathé, François Remy +3
While Large Language Models (LLMs) have shown remarkable capabilities in natural language understanding and generation, their performance often lags in lower-resource, non-English…
Achilles' Heels: Vulnerable Record Identification in Synthetic Data Publishing
Matthieu Meeus, Florent Guépin, Ana-Maria Cretu +1
Synthetic data is seen as the most promising solution to share individual-level data while preserving privacy. Shadow modeling-based Membership Inference Attacks (MIAs) have become…
The Canary's Echo: Auditing Privacy Risks of LLM-Generated Synthetic Text
Matthieu Meeus, Lukas Wutschitz, Santiago Zanella-Béguelin +2
How much information about training samples can be leaked through synthetic data generated by Large Language Models (LLMs)? Overlooking the subtleties of information flow in synthe…
Checkpoint-GCG: Auditing and Attacking Fine-Tuning-Based Prompt Injection Defenses
Xiaoxue Yang, Bozhidar Stevanoski, Matthieu Meeus +1
Large language models (LLMs) are increasingly deployed in real-world applications ranging from chatbots to agentic systems, where they are expected to process untrusted data and fo…
Synthetic is all you need: removing the auxiliary data assumption for membership inference attacks against synthetic data
Florent Guépin, Matthieu Meeus, Ana-Maria Cretu +1
Synthetic data is emerging as one of the most promising solutions to share individual-level data while safeguarding privacy. While membership inference attacks (MIAs), based on sha…