7 papers · 1 filter
RAT-Bench: A Comprehensive Benchmark for Text Anonymization
NataÅ¡a KrÄo, Zexi Yao, Matthieu Meeus +1
Data containing personal information is increasingly used to train, fine-tune, or query Large Language Models (LLMs). Text is typically scrubbed of identifying information prior to…
The Canary's Echo: Auditing Privacy Risks of LLM-Generated Synthetic Text
Matthieu Meeus, Lukas Wutschitz, Santiago Zanella-Béguelin +2
How much information about training samples can be leaked through synthetic data generated by Large Language Models (LLMs)? Overlooking the subtleties of information flow in synthe…
The Mosaic Memory of Large Language Models
Igor Shilov, Matthieu Meeus, Yves-Alexandre de Montjoye
As Large Language Models (LLMs) become widely adopted, understanding how they learn from, and memorize, training data becomes crucial. Memorization in LLMs is widely assumed to onl…
SoK: Membership Inference Attacks on LLMs are Rushing Nowhere (and How to Fix It)
Matthieu Meeus, Igor Shilov, Shubham Jain +3
Whether LLMs memorize their training data and what this means, from measuring privacy leakage to detecting copyright violations, has become a rapidly growing area of research. In t…
ChocoLlama: Lessons Learned From Teaching Llamas Dutch
Matthieu Meeus, Anthony Rathé, François Remy +3
While Large Language Models (LLMs) have shown remarkable capabilities in natural language understanding and generation, their performance often lags in lower-resource, non-English…
Did the Neurons Read your Book? Document-level Membership Inference for Large Language Models
Matthieu Meeus, Shubham Jain, Marek Rei +1
With large language models (LLMs) poised to become embedded in our daily lives, questions are starting to be raised about the data they learned from. These questions range from pot…