Benchmarking Differentially Private Synthetic Data Generation Algorithms
arXiv:2112.09238
Abstract
This work presents a systematic benchmark of differentially private synthetic data generation algorithms that can generate tabular data. Utility of the synthetic data is evaluated by measuring whether the synthetic data preserve the distribution of individual and pairs of attributes, pairwise correlation as well as on the accuracy of an ML classification model. In a comprehensive empirical evaluation we identify the top performing algorithms and those that consistently fail to beat baseline approaches.
Cited by in corpus (5)
- How to DP-fy ML: A Practical Guide to Machine Learning with Differential Privacy
- Synthetic Data: Revisiting the Privacy-Utility Trade-off
- DPVisCreator: Incorporating Pattern Constraints to Privacy-preserving Visualizations via Differential Privacy
- Synthetic Tabular Data: Methods, Attacks and Defenses
- Detecting Anomalous LAN Activities under Differential Privacy