10 papers
Detecting Functional Memorization in Code Language Models
Matthieu Meeus, Anil Ramakrishna, Shengyuan Hu +3
Large language models (LLMs) are increasingly used to generate code at scale. Meanwhile, prior work has investigated whether training data may be recoverable from model outputs, by…
RAT-Bench: A Comprehensive Benchmark for Text Anonymization
NataÅ¡a KrÄo, Zexi Yao, Matthieu Meeus +1
Data containing personal information is increasingly used to train, fine-tune, or query Large Language Models (LLMs). Text is typically scrubbed of identifying information prior to…
Exploring the limits of strong membership inference attacks on large language models
Jamie Hayes, Ilia Shumailov, Christopher A. Choquette-Choo +13
State-of-the-art membership inference attacks (MIAs) typically require training many reference models, making it difficult to scale these attacks to large pre-trained language mode…
Checkpoint-GCG: Auditing and Attacking Fine-Tuning-Based Prompt Injection Defenses
Xiaoxue Yang, Bozhidar Stevanoski, Matthieu Meeus +1
Large language models (LLMs) are increasingly deployed in real-world applications ranging from chatbots to agentic systems, where they are expected to process untrusted data and fo…
Lost in the Averages: A New Specific Setup to Evaluate Membership Inference Attacks Against Machine Learning Models
NataÅ¡a KrÄo, Florent Guépin, Matthieu Meeus +2
Synthetic data generators and machine learning models can memorize their training data, posing privacy concerns. Membership inference attacks (MIAs) are a standard method of estima…
Counterfactual Influence as a Distributional Quantity
Matthieu Meeus, Igor Shilov, Georgios Kaissis +1
Machine learning models are known to memorize samples from their training data, raising concerns around privacy and generalization. Counterfactual self-influence is a popular metri…