5 papers · 1 filter
Rethinking Accuracy: A Weighted Error-Based Metric for Data Quality
Valerie Restat, Uta Störl
Real data often contains errors, which is why data engineers spend a lot of time creating data cleaning pipelines to ensure the best possible data quality. However, it is often dif…
Extending GouDa: Generation of Universal Datasets with (and without) Errors for Data Quality Benchmarking
Valerie Restat, André Conrad, Kevin M. Kramer +1
Synthetic data is extremely important in areas such as data quality, data cleaning, and machine learning. It enables the analysis of use cases in which real data is insufficient, u…
Towards Next Generation Data Engineering Pipelines
Kevin M. Kramer, Valerie Restat, Sebastian Strasser +2
Data engineering pipelines are a widespread way to provide high-quality data for all kinds of data science applications. However, numerous challenges still remain in the compositio…
Data Cleaning of Data Streams
Valerie Restat, Niklas Rodenhausen, Carina Antonin +1
Streaming data can arise from a variety of contexts. Important use cases are continuous sensor measurements such as temperature, light or radiation values. In the process, streamin…
MVIAnalyzer: A Holistic Approach to Analyze Missing Value Imputation
Valerie Restat, Kai Tejkl, Uta Störl
Missing values often limit the usage of data analysis or cause falsification of results. Therefore, methods of missing value imputation (MVI) are of great significance. However, in…