Showing cs.DBShow all
2 papers · 1 filter
cs.DB2026
Rethinking Accuracy: A Weighted Error-Based Metric for Data Quality
Valerie Restat, Uta Störl
Real data often contains errors, which is why data engineers spend a lot of time creating data cleaning pipelines to ensure the best possible data quality. However, it is often dif…
cs.DB2026
Extending GouDa: Generation of Universal Datasets with (and without) Errors for Data Quality Benchmarking
Valerie Restat, André Conrad, Kevin M. Kramer +1
Synthetic data is extremely important in areas such as data quality, data cleaning, and machine learning. It enables the analysis of use cases in which real data is insufficient, u…