10 papers
How to DP-fy Your Data: A Practical Guide to Generating Synthetic Data With Differential Privacy
Natalia Ponomareva, Zheng Xu, H. Brendan McMahan +12
High quality data is needed to unlock the full potential of AI for end users. However finding new sources of such data is getting harder: most publicly-available human generated da…
Advancing the State-of-the-Art in Empirical Privacy Auditing
Nicole Mitchell, Galen Andrew, Arun Ganesh +2
Parameter-efficient fine-tuning of large language models (LLMs) can exhibit problematic memorization of individual training examples. Empirical privacy auditing (EPA) quantifies th…
CLIOPATRA: Extracting Private Information from LLM Insights
Meenatchi Sundaram Muthu Selva Annamalai, Emiliano De Cristofaro, Peter Kairouz
The widespread adoption of AI assistants has prompted the development of privacy-aware platforms designed to extract insights from real-world usage. Their privacy protections prima…
ACTG-ARL: Differentially Private Conditional Text Generation with RL-Boosted Control
Yuzheng Hu, Ryan McKenna, Da Yu +4
Generating high-quality synthetic text under differential privacy (DP) is critical for training and evaluating language models without compromising user privacy. Prior work on synt…
Mayfly: Private Aggregate Insights from Ephemeral Streams of On-Device User Data
Christopher Bian, Albert Cheu, Stanislav Chiknavaryan +12
This paper introduces Mayfly, a federated analytics approach enabling aggregate queries over ephemeral on-device data streams without central persistence of sensitive user data. Ma…
Toward provably private analytics and insights into GenAI use
Albert Cheu, Artem Lagzdin, Brett McLarnon +8
Large-scale systems that compute analytics over a fleet of devices must achieve high privacy and security standards while also meeting data quality, usability, and resource efficie…