activity
20242026
collaborators

6 papers

cs.LG2026

Phantoms and Disclosures: A Statistical Framework for Auditing Privacy in Synthetic Data

Kareem Amin, Rudrajit Das, Alessandro Epasto +4

The rapid adoption of generative AI and Large Language Models (LLMs) has spurred interest in synthetic data as a privacy-preserving alternative to sensitive real-world datasets. Ho…

cs.LG2026

Learning from Synthetic Data: Limitations of ERM

Kareem Amin, Alex Bie, Weiwei Kong +2

The prevalence and low cost of LLMs have led to a rise of synthetic content. From review sites to court documents, "natural" content has been contaminated by data points that appea…

cs.LG2025

Escaping Collapse: The Strength of Weak Data for Large Language Model Training

Kareem Amin, Sara Babakniya, Alex Bie +3

Synthetically-generated data plays an increasingly larger role in training large language models. However, while synthetic data has been found to be useful, studies have also shown…

cs.LG2025

Clustering and Median Aggregation Improve Differentially Private Inference

Kareem Amin, Salman Avestimehr, Sara Babakniya +4

Differentially private (DP) language model inference is an approach for generating private synthetic text. A sensitive input example is used to prompt an off-the-shelf large langua…

cs.LG2024

Private prediction for large-scale synthetic text generation

Kareem Amin, Alex Bie, Weiwei Kong +5

We present an approach for generating differentially private synthetic text using large language models (LLMs), via private prediction. In the private prediction framework, we only…

cs.CR2024

Practical Considerations for Differential Privacy

Kareem Amin, Alex Kulesza, Sergei Vassilvitskii

Differential privacy is the gold standard for statistical data release. Used by governments, companies, and academics, its mathematically rigorous guarantees and worst-case assumpt…