A Consensus Privacy Metrics Framework for Synthetic Data
arXiv:2503.04980 · doi:10.1016/j.patter.2025.101320
Abstract
Synthetic data generation is one approach for sharing individual-level data. However, to meet legislative requirements, it is necessary to demonstrate that the individuals' privacy is adequately protected. There is no consolidated standard for measuring privacy in synthetic data. Through an expert panel and consensus process, we developed a framework for evaluating privacy in synthetic data. Our findings indicate that current similarity metrics fail to measure identity disclosure, and their use is discouraged. For differentially private synthetic data, a privacy budget other than close to zero was not considered interpretable. There was consensus on the importance of membership and attribute disclosure, both of which involve inferring personal information about an individual without necessarily revealing their identity. The resultant framework provides precise recommendations for metrics that address these types of disclosures effectively. Our findings further present specific opportunities for future research that can help with widespread adoption of synthetic data.
References in corpus (16)
- Data Synthesis based on Generative Adversarial Networks
- A Multifaceted Benchmarking of Synthetic Electronic Health Record Generation Models
- A review of Generative Adversarial Networks for Electronic Health Records: applications, evaluation measures and data sources
- Can I trust my fake data -- A comprehensive quality assessment framework for synthetic tabular data in healthcare
- Membership Inference Attacks against Synthetic Data through Overfitting Detection
- Differentially Private Synthetic Data: Applied Evaluations and Enhancements
- Fidelity and Privacy of Synthetic Medical Data
- TAPAS: a Toolbox for Adversarial Privacy Auditing of Synthetic Data
- Performance evaluation of predictive AI models to support medical decisions: Overview and guidance
- Differentially Private Release of Israel's National Registry of Live Births
- Privacy Measurement in Tabular Synthetic Data: State of the Art and Future Research Directions
- A Members First Approach to Enabling LinkedIn's Labor Market Insights at Scale
- Synthetic Data and Health Privacy
- Selecting a classification performance measure: matching the measure to the problem
- Towards more accurate and useful data anonymity vulnerability measures
- Practical privacy metrics for synthetic data