7 papers
Phantoms and Disclosures: A Statistical Framework for Auditing Privacy in Synthetic Data
Kareem Amin, Rudrajit Das, Alessandro Epasto +4
The rapid adoption of generative AI and Large Language Models (LLMs) has spurred interest in synthetic data as a privacy-preserving alternative to sensitive real-world datasets. Ho…
Is your algorithm unlearning or untraining?
Eleni Triantafillou, Ahmed Imtiaz Humayun, Monica Ribero +3
As models are getting larger and are trained on increasing amounts of data, there has been an explosion of interest into how we can ``delete'' specific data points or behaviours fr…
Regularized -Divergence Kernel Tests
Mónica Ribero, Antonin Schrab, Arthur Gretton
We propose a framework to construct practical kernel-based two-sample tests from the family of -divergences. The test statistic is computed from the witness function of a regula…
Sequentially Auditing Differential Privacy
Tomás González, Mateo Dulce-Rubio, Aaditya Ramdas +1
We propose a practical sequential test for auditing differential privacy guarantees of black-box mechanisms. The test processes streams of mechanisms' outputs providing anytime-val…
Differentially Private Optimization for Non-Decomposable Objective Functions
Weiwei Kong, Andrés Muñoz Medina, Mónica Ribero
Unsupervised pre-training is a common step in developing computer vision models and large language models. In this setting, the absence of labels requires the use of similarity-bas…
Privacy of the last iterate in cyclically-sampled DP-SGD on nonconvex composite losses
Weiwei Kong, Mónica Ribero
Differentially-private stochastic gradient descent (DP-SGD) is a family of iterative machine learning training algorithms that privatize gradients to generate a sequence of differe…