activity
20242026
collaborators

6 papers

cs.DB2026

A Catalog of Data Errors

Divya Bhadauria, Hazar Harmouch, Felix Naumann +2

Data errors are widespread in real-world databases and severely impact downstream applications, such as machine learning pipelines or business analytics reports. Causes of such err…

cs.DB2025

Is SHACL Suitable for Data Quality Assessment?

Carolina Cortés, Lisa Ehrlinger, Lorena Etcheverry +1

Knowledge graphs have been widely adopted in both enterprises, such as the Google Knowledge Graph, and open platforms like Wikidata, to represent domain knowledge and support artif…

cs.DB2025

Enabling Data Dependency-based Query Optimization

Daniel Lindner, Daniel Ritter, Felix Naumann

Primary key (PK) and foreign key (FK) constraints are widely used for query optimization. Knowledge about additional data dependencies, such as order dependencies, enables further…

cs.DB2025

The Effects of Data Quality on Machine Learning Performance on Tabular Data

Sedir Mohammed, Lukas Budach, Moritz Feuerpfeil +6

Modern artificial intelligence (AI) applications require large quantities of training and test data. This need creates critical challenges not only concerning the availability of s…

cs.DB2025

Step-by-Step Data Cleaning Recommendations to Improve ML Prediction Accuracy

Sedir Mohammed, Felix Naumann, Hazar Harmouch

Data quality is crucial in machine learning (ML) applications, as errors in the data can significantly impact the prediction accuracy of the underlying ML model. Therefore, data cl…

cs.DB2024

Data Quality Assessment: Challenges and Opportunities

Sedir Mohammed, Lisa Ehrlinger, Hazar Harmouch +2

Data-oriented applications, their users, and even the law require data of high quality. Research has divided the rather vague notion of data quality into various dimensions, such a…