6 papers
SemPiper: Interactive Code Synthesis for Semantic Operators in Machine Learning Pipelines
Olga Ovcharenko, Luciano Duarte, Sebastian Schelter
Machine learning (ML) pipelines require extensive data preparation, feature engineering, and integration across heterogeneous sources, making them tedious and error-prone to develo…
ArtiFact: A Large-Scale Multi-Modal Cultural Heritage Dataset
Luciano Duarte, Olga Ovcharenko, Sebastian Schelter
Multi-modal data management has emerged as a central research topic in the database community, spanning data integration, semantic query processing, and data quality assessment. De…
SemBench: A Benchmark for Semantic Query Processing Engines
Jiale Lao, Andreas Zimmerer, Olga Ovcharenko +12
We present a benchmark targeting a novel class of systems: semantic query processing engines. Those systems rely inherently on generative and reasoning capabilities of state-of-the…
SemPipes -- Optimizable Semantic Data Operators for Tabular Machine Learning Pipelines
Olga Ovcharenko, Matthias Boehm, Sebastian Schelter
Real-world machine learning on tabular data relies on complex data preparation pipelines for prediction, data integration, augmentation, and debugging. Designing these pipelines re…
Towards Cross-Modal Error Detection with Tables and Images
Olga Ovcharenko, Sebastian Schelter
Ensuring data quality at scale remains a persistent challenge for large organizations. Despite recent advances, maintaining accurate and consistent data is still complex, especiall…
Towards a Real-World Aligned Benchmark for Unlearning in Recommender Systems
Pierre Lubitzsch, Olga Ovcharenko, Hao Chen +2
Modern recommender systems heavily leverage user interaction data to deliver personalized experiences. However, relying on personal data presents challenges in adhering to privacy…