4 papers
Privacy-Preserving Tabular Synthetic Data Generation Using TabularARGN
Andrey Sidorenko, Paul Tiwald
Synthetic data generation has become essential for securely sharing and analyzing sensitive data sets. Traditional anonymization techniques, however, often fail to adequately prese…
Improving Predictions on Highly Unbalanced Data Using Open Source Synthetic Data Upsampling
Ivona Krchova, Michael Platzer, Paul Tiwald
Unbalanced tabular data sets present significant challenges for predictive modeling and data analysis across a wide range of applications. In many real-world scenarios, such as fra…
Benchmarking Synthetic Tabular Data: A Multi-Dimensional Evaluation Framework
Andrey Sidorenko, Michael Platzer, Mario Scriminaci +1
Evaluating the quality of synthetic data remains a key challenge for ensuring privacy and utility in data-driven research. In this work, we present an evaluation framework that qua…
TabularARGN: A Flexible and Efficient Auto-Regressive Framework for Generating High-Fidelity Synthetic Data
Paul Tiwald, Ivona Krchova, Andrey Sidorenko +3
Synthetic data generation for tabular datasets must balance fidelity, efficiency, and versatility to meet the demands of real-world applications. We introduce the Tabular Auto-Regr…