activity
20242026
most citedScaling Laws for Downstream Task Performance of Large Language Models

3 citations · 5 across the 9 of their papers we have counts for

collaborators

10 papers

cs.LG2026

Phantoms and Disclosures: A Statistical Framework for Auditing Privacy in Synthetic Data

Kareem Amin, Rudrajit Das, Alessandro Epasto +4

The rapid adoption of generative AI and Large Language Models (LLMs) has spurred interest in synthetic data as a privacy-preserving alternative to sensitive real-world datasets. Ho…

cs.LG2026

AI-rithmetic

Alex Bie, Travis Dick, Alex Kulesza +3

Modern AI systems have been successfully deployed to win medals at international math competitions, assist with research workflows, and prove novel technical lemmas. However, despi…

cs.LG2026

Learning from Synthetic Data: Limitations of ERM

Kareem Amin, Alex Bie, Weiwei Kong +2

The prevalence and low cost of LLMs have led to a rise of synthetic content. From review sites to court documents, "natural" content has been contaminated by data points that appea…

cs.CR2025★ 1 cited

How to DP-fy Your Data: A Practical Guide to Generating Synthetic Data With Differential Privacy

Natalia Ponomareva, Zheng Xu, H. Brendan McMahan +12

High quality data is needed to unlock the full potential of AI for end users. However finding new sources of such data is getting harder: most publicly-available human generated da…

cs.CR2025

Differentially Private Synthetic Data Release for Topics API Outputs

Travis Dick, Alessandro Epasto, Adel Javanmard +5

The analysis of the privacy properties of Privacy-Preserving Ads APIs is an area of research that has received strong interest from academics, industry, and regulators. Despite thi…

cs.LG2025

An Optimization Framework for Differentially Private Sparse Fine-Tuning

Mehdi Makni, Kayhan Behdin, Gabriel Afriat +5

Differentially private stochastic gradient descent (DP-SGD) is broadly considered to be the gold standard for training and fine-tuning neural networks under differential privacy (D…