3 papers
cs.LG2025
Democratizing Tabular Data Access with an Open$\unicode{x2013}$Source Synthetic$\unicode{x2013}$Data SDK
Ivona Krchova, Mariana Vargas Vieyra, Mario Scriminaci +1
Machine learning development critically depends on access to high-quality data. However, increasing restrictions due to privacy, proprietary interests, and ethical concerns have cr…
cs.LG2025
Improving Predictions on Highly Unbalanced Data Using Open Source Synthetic Data Upsampling
Ivona Krchova, Michael Platzer, Paul Tiwald
Unbalanced tabular data sets present significant challenges for predictive modeling and data analysis across a wide range of applications. In many real-world scenarios, such as fra…
cs.LG2025
TabularARGN: A Flexible and Efficient Auto-Regressive Framework for Generating High-Fidelity Synthetic Data
Paul Tiwald, Ivona Krchova, Andrey Sidorenko +3
Synthetic data generation for tabular datasets must balance fidelity, efficiency, and versatility to meet the demands of real-world applications. We introduce the Tabular Auto-Regr…