5 papers
Democratizing Tabular Data Access with an Open$\unicode{x2013}$Source Synthetic$\unicode{x2013}$Data SDK
Ivona Krchova, Mariana Vargas Vieyra, Mario Scriminaci +1
Machine learning development critically depends on access to high-quality data. However, increasing restrictions due to privacy, proprietary interests, and ethical concerns have cr…
Privacy-Preserving Tabular Synthetic Data Generation Using TabularARGN
Andrey Sidorenko, Paul Tiwald
Synthetic data generation has become essential for securely sharing and analyzing sensitive data sets. Traditional anonymization techniques, however, often fail to adequately prese…
A Note on Statistically Accurate Tabular Data Generation Using Large Language Models
Andrey Sidorenko
Large language models (LLMs) have shown promise in synthetic tabular data generation, yet existing methods struggle to preserve complex feature dependencies, particularly among cat…
Benchmarking Synthetic Tabular Data: A Multi-Dimensional Evaluation Framework
Andrey Sidorenko, Michael Platzer, Mario Scriminaci +1
Evaluating the quality of synthetic data remains a key challenge for ensuring privacy and utility in data-driven research. In this work, we present an evaluation framework that qua…
TabularARGN: A Flexible and Efficient Auto-Regressive Framework for Generating High-Fidelity Synthetic Data
Paul Tiwald, Ivona Krchova, Andrey Sidorenko +3
Synthetic data generation for tabular datasets must balance fidelity, efficiency, and versatility to meet the demands of real-world applications. We introduce the Tabular Auto-Regr…