6 papers
Phantoms and Disclosures: A Statistical Framework for Auditing Privacy in Synthetic Data
Kareem Amin, Rudrajit Das, Alessandro Epasto +4
The rapid adoption of generative AI and Large Language Models (LLMs) has spurred interest in synthetic data as a privacy-preserving alternative to sensitive real-world datasets. Ho…
Learning from Synthetic Data: Limitations of ERM
Kareem Amin, Alex Bie, Weiwei Kong +2
The prevalence and low cost of LLMs have led to a rise of synthetic content. From review sites to court documents, "natural" content has been contaminated by data points that appea…
Escaping Collapse: The Strength of Weak Data for Large Language Model Training
Kareem Amin, Sara Babakniya, Alex Bie +3
Synthetically-generated data plays an increasingly larger role in training large language models. However, while synthetic data has been found to be useful, studies have also shown…
Clustering and Median Aggregation Improve Differentially Private Inference
Kareem Amin, Salman Avestimehr, Sara Babakniya +4
Differentially private (DP) language model inference is an approach for generating private synthetic text. A sensitive input example is used to prompt an off-the-shelf large langua…
Private prediction for large-scale synthetic text generation
Kareem Amin, Alex Bie, Weiwei Kong +5
We present an approach for generating differentially private synthetic text using large language models (LLMs), via private prediction. In the private prediction framework, we only…
Practical Considerations for Differential Privacy
Kareem Amin, Alex Kulesza, Sergei Vassilvitskii
Differential privacy is the gold standard for statistical data release. Used by governments, companies, and academics, its mathematically rigorous guarantees and worst-case assumpt…