most citedHow to DP-fy Your Data: A Practical Guide to Generating Synthetic Data With Differential Privacy

1 citations · 1 across the 2 of their papers we have counts for

collaborators

5 papers

cs.CR20261 cited

How to DP-fy Your Data: A Practical Guide to Generating Synthetic Data With Differential Privacy

Natalia Ponomareva, Zheng Xu, H. Brendan McMahan +12

High quality data is needed to unlock the full potential of AI for end users. However finding new sources of such data is getting harder: most publicly-available human generated da…

cs.LG2026

ContinuousBench: Can Differentially Private Synthetic Text Improve Capabilities?

Peihan Liu, Lucas Rosenblatt, Weiwei Kong +7

Differentially private (DP) text synthesis promises to unlock sensitive corpora for model training, but it remains unclear whether DP synthetic data transmits genuinely new knowled…

cs.LG2026

Learning from Synthetic Data: Limitations of ERM

Kareem Amin, Alex Bie, Weiwei Kong +2

The prevalence and low cost of LLMs have led to a rise of synthetic content. From review sites to court documents, "natural" content has been contaminated by data points that appea…

cs.LG2025

Escaping Collapse: The Strength of Weak Data for Large Language Model Training

Kareem Amin, Sara Babakniya, Alex Bie +3

Synthetically-generated data plays an increasingly larger role in training large language models. However, while synthetic data has been found to be useful, studies have also shown…

cs.LG2025

Clustering and Median Aggregation Improve Differentially Private Inference

Kareem Amin, Salman Avestimehr, Sara Babakniya +4

Differentially private (DP) language model inference is an approach for generating private synthetic text. A sensitive input example is used to prompt an off-the-shelf large langua…