7 papers
Rethinking Accuracy: A Weighted Error-Based Metric for Data Quality
Valerie Restat, Uta Störl
Real data often contains errors, which is why data engineers spend a lot of time creating data cleaning pipelines to ensure the best possible data quality. However, it is often dif…
Extending GouDa: Generation of Universal Datasets with (and without) Errors for Data Quality Benchmarking
Valerie Restat, André Conrad, Kevin M. Kramer +1
Synthetic data is extremely important in areas such as data quality, data cleaning, and machine learning. It enables the analysis of use cases in which real data is insufficient, u…
Solving Distributed Flexible Job Shop Scheduling Problems in the Wool Textile Industry with Quantum Annealing
Lilia Toma, Markus Zajac, Uta Störl
Many modern manufacturing companies have evolved from a single production facility to a multi-factory production environment that must manage both regionally dispersed production o…
QC-Adviser: Quantum Hardware Recommendations for Solving Industrial Optimization Problems
Djamel Laps-Bouraba, Markus Zajac, Uta Störl
The availability of quantum hardware via the cloud offers opportunities for new approaches to computing optimization problems in an industrial environment. However, selecting the r…
Towards Next Generation Data Engineering Pipelines
Kevin M. Kramer, Valerie Restat, Sebastian Strasser +2
Data engineering pipelines are a widespread way to provide high-quality data for all kinds of data science applications. However, numerous challenges still remain in the compositio…
Data Cleaning of Data Streams
Valerie Restat, Niklas Rodenhausen, Carina Antonin +1
Streaming data can arise from a variety of contexts. Important use cases are continuous sensor measurements such as temperature, light or radiation values. In the process, streamin…