collaborators
Showing cs.DBShow all

5 papers · 1 filter

cs.DB2026

Rethinking Accuracy: A Weighted Error-Based Metric for Data Quality

Valerie Restat, Uta Störl

Real data often contains errors, which is why data engineers spend a lot of time creating data cleaning pipelines to ensure the best possible data quality. However, it is often dif…

cs.DB2026

Extending GouDa: Generation of Universal Datasets with (and without) Errors for Data Quality Benchmarking

Valerie Restat, André Conrad, Kevin M. Kramer +1

Synthetic data is extremely important in areas such as data quality, data cleaning, and machine learning. It enables the analysis of use cases in which real data is insufficient, u…

cs.DB2025

Towards Next Generation Data Engineering Pipelines

Kevin M. Kramer, Valerie Restat, Sebastian Strasser +2

Data engineering pipelines are a widespread way to provide high-quality data for all kinds of data science applications. However, numerous challenges still remain in the compositio…

cs.DB2025

Data Cleaning of Data Streams

Valerie Restat, Niklas Rodenhausen, Carina Antonin +1

Streaming data can arise from a variety of contexts. Important use cases are continuous sensor measurements such as temperature, light or radiation values. In the process, streamin…

cs.DB2025

MVIAnalyzer: A Holistic Approach to Analyze Missing Value Imputation

Valerie Restat, Kai Tejkl, Uta Störl

Missing values often limit the usage of data analysis or cause falsification of results. Therefore, methods of missing value imputation (MVI) are of great significance. However, in…