Showing cs.DBShow all
2 papers · 1 filter
cs.DB2026
Extending GouDa: Generation of Universal Datasets with (and without) Errors for Data Quality Benchmarking
Valerie Restat, André Conrad, Kevin M. Kramer +1
Synthetic data is extremely important in areas such as data quality, data cleaning, and machine learning. It enables the analysis of use cases in which real data is insufficient, u…
cs.DB2025
Towards Next Generation Data Engineering Pipelines
Kevin M. Kramer, Valerie Restat, Sebastian Strasser +2
Data engineering pipelines are a widespread way to provide high-quality data for all kinds of data science applications. However, numerous challenges still remain in the compositio…