collaborators

13 papers

cs.DB2026

"Will This Data Break My Task?" - Interactive Synthesis of Task-Aware Data Unit Tests

Hao Chen, Arnab Phani, Sebastian Schelter

Data is a central resource for modern enterprises and institutions, and data errors propagating through data pipelines lead to serious impact in production. Therefore, data validat…

cs.DB2026

Beyond Scale and Generation: Understanding Language Model-based Entity Matching

Zeyu Zhang, Xue Li, Iacer Calixto +2

Entity matching identifies records that refer to the same real-world entity. Language models can be adapted to this task through bi-encoder, cross-encoder, and generative matcher a…

cs.LG2026

SemPiper: Interactive Code Synthesis for Semantic Operators in Machine Learning Pipelines

Olga Ovcharenko, Luciano Duarte, Sebastian Schelter

Machine learning (ML) pipelines require extensive data preparation, feature engineering, and integration across heterogeneous sources, making them tedious and error-prone to develo…

cs.DB2026

ArtiFact: A Large-Scale Multi-Modal Cultural Heritage Dataset

Luciano Duarte, Olga Ovcharenko, Sebastian Schelter

Multi-modal data management has emerged as a central research topic in the database community, spanning data integration, semantic query processing, and data quality assessment. De…

cs.LG2026

Be Fair! Can Machine Learning Engineering Agents Adhere to Fairness Constraints?

Anna Richter, Julia Stoyanovich, Sebastian Schelter

Machine learning engineering (MLE) agents promise to automate end-to-end ML pipeline development from raw data and natural language instructions, potentially making ML accessible t…

cs.LG2026

PrismaDV: Automated Task-Aware Data Unit Test Generation

Hao Chen, Arnab Phani, Sebastian Schelter

Data is a central resource for modern enterprises, and data validation is essential for ensuring the reliability of downstream applications. However, existing automated data unit t…