statistics

tidysynthesis: a Meta-Package for Synthetic Data Generation

arXiv:2607.12611

summary

The paper presents tidysynthesis, a meta‑package that provides a unified syntax for building and iterating synthetic data generation pipelines, improving interoperability between modeling tools and statistical privacy methods.

Abstract

Synthetic data generation enables data curators to more easily share datasets that limits the potential for disclosive inferences about data subjects in confidential datasets. Generating synthetic data requires navigating numerous design choices; however, most existing open source software fails to provide common software infrastructure for making such design choices efficiently. In this paper, we introduce tidysynthesis, a meta-package for synthetic data generation that enables better interoperability between existing modeling frameworks and statistical data privacy methods. tidysynthesis allows users more flexibility to specify and iterate on synthetic data algorithms by providing a common syntax to easily create and modify synthetic data generation pipelines. We demonstrate the features and extensibility of tidysynthesis, as well as provide end-to-end examples for synthetic data generation using data from the American Community Survey

23 pages

Topics & keywords

#synthetic data#data privacy#software infrastructure#pipeline design#interoperability#statistical disclosure controlsynthetic data generationmeta-packageprivacy-preserving datapipelineAmerican Community Surveystatistical disclosure control
tidysynthesis: a Meta-Package for Synthetic Data Generation · wovepaper