collaborators

6 papers

cs.LG2026

ContinuousBench: Can Differentially Private Synthetic Text Improve Capabilities?

Peihan Liu, Lucas Rosenblatt, Weiwei Kong +7

Differentially private (DP) text synthesis promises to unlock sensitive corpora for model training, but it remains unclear whether DP synthetic data transmits genuinely new knowled…

cs.CV2026

Decomposing Private Image Generation via Coarse-to-Fine Wavelet Modeling

Jasmine Bayrooti, Weiwei Kong, Natalia Ponomareva +3

Generative models trained on sensitive image datasets risk memorizing and reproducing individual training examples, making strong privacy guarantees essential. While differential p…

cs.LG2026

JAX-Privacy: A library for differentially private machine learning

Ryan McKenna, Galen Andrew, Borja Balle +6

JAX-Privacy is a library designed to simplify the deployment of robust and performant mechanisms for differentially private machine learning. Guided by design principles of usabili…

cs.LG2026

Learning from Synthetic Data: Limitations of ERM

Kareem Amin, Alex Bie, Weiwei Kong +2

The prevalence and low cost of LLMs have led to a rise of synthetic content. From review sites to court documents, "natural" content has been contaminated by data points that appea…

cs.LG2025

Escaping Collapse: The Strength of Weak Data for Large Language Model Training

Kareem Amin, Sara Babakniya, Alex Bie +3

Synthetically-generated data plays an increasingly larger role in training large language models. However, while synthetic data has been found to be useful, studies have also shown…

cs.LG2025

Clustering and Median Aggregation Improve Differentially Private Inference

Kareem Amin, Salman Avestimehr, Sara Babakniya +4

Differentially private (DP) language model inference is an approach for generating private synthetic text. A sensitive input example is used to prompt an off-the-shelf large langua…