13 papers
SMOTE and Mirrors: Exposing Privacy Leakage from Synthetic Minority Oversampling
Georgi Ganev, Reza Nazari, Rees Davison +5
The Synthetic Minority Over-sampling Technique (SMOTE) is one of the most widely used methods for addressing class imbalance and generating synthetic data. Despite its popularity,…
Rethinking Anonymity Claims in Synthetic Data Generation: A Model-Centric Privacy Attack Perspective
Georgi Ganev, Emiliano De Cristofaro
Training generative machine learning models to produce synthetic tabular data has become a popular approach for enhancing privacy in data sharing. As this typically involves proces…
The Hitchhiker's Guide to Efficient, End-to-End, and Tight DP Auditing
Meenatchi Sundaram Muthu Selva Annamalai, Borja Balle, Jamie Hayes +2
In this paper, we systematize research on auditing Differential Privacy (DP) techniques, aiming to identify key insights and open challenges. First, we introduce a comprehensive fr…
To Shuffle or not to Shuffle: Auditing DP-SGD with Shuffling
Meenatchi Sundaram Muthu Selva Annamalai, Borja Balle, Jamie Hayes +1
The Differentially Private Stochastic Gradient Descent (DP-SGD) algorithm supports the training of machine learning (ML) models with formal Differential Privacy (DP) guarantees. Tr…
The Importance of Being Discrete: Measuring the Impact of Discretization in End-to-End Differentially Private Synthetic Data
Georgi Ganev, Meenatchi Sundaram Muthu Selva Annamalai, Sofiane Mahiou +1
Differentially Private (DP) generative marginal models are often used in the wild to release synthetic tabular datasets in lieu of sensitive data while providing formal privacy gua…
The Inadequacy of Similarity-based Privacy Metrics: Privacy Attacks against "Truly Anonymous" Synthetic Datasets
Georgi Ganev, Emiliano De Cristofaro
Generative models producing synthetic data are meant to provide a privacy-friendly approach to releasing data. However, their privacy guarantees are only considered robust when mod…