2 papers
cs.CL2026
RAT-Bench: A Comprehensive Benchmark for Text Anonymization
NataÅ¡a KrÄo, Zexi Yao, Matthieu Meeus +1
Data containing personal information is increasingly used to train, fine-tune, or query Large Language Models (LLMs). Text is typically scrubbed of identifying information prior to…
cs.CR2025
The DCR Delusion: Measuring the Privacy Risk of Synthetic Data
Zexi Yao, NataÅ¡a KrÄo, Georgi Ganev +1
Synthetic data has become an increasingly popular way to share data without revealing sensitive information. Though Membership Inference Attacks (MIAs) are widely considered the go…