13 papers
"Will This Data Break My Task?" - Interactive Synthesis of Task-Aware Data Unit Tests
Hao Chen, Arnab Phani, Sebastian Schelter
Data is a central resource for modern enterprises and institutions, and data errors propagating through data pipelines lead to serious impact in production. Therefore, data validat…
Beyond Scale and Generation: Understanding Language Model-based Entity Matching
Zeyu Zhang, Xue Li, Iacer Calixto +2
Entity matching identifies records that refer to the same real-world entity. Language models can be adapted to this task through bi-encoder, cross-encoder, and generative matcher a…
SemPiper: Interactive Code Synthesis for Semantic Operators in Machine Learning Pipelines
Olga Ovcharenko, Luciano Duarte, Sebastian Schelter
Machine learning (ML) pipelines require extensive data preparation, feature engineering, and integration across heterogeneous sources, making them tedious and error-prone to develo…
ArtiFact: A Large-Scale Multi-Modal Cultural Heritage Dataset
Luciano Duarte, Olga Ovcharenko, Sebastian Schelter
Multi-modal data management has emerged as a central research topic in the database community, spanning data integration, semantic query processing, and data quality assessment. De…
Be Fair! Can Machine Learning Engineering Agents Adhere to Fairness Constraints?
Anna Richter, Julia Stoyanovich, Sebastian Schelter
Machine learning engineering (MLE) agents promise to automate end-to-end ML pipeline development from raw data and natural language instructions, potentially making ML accessible t…
PrismaDV: Automated Task-Aware Data Unit Test Generation
Hao Chen, Arnab Phani, Sebastian Schelter
Data is a central resource for modern enterprises, and data validation is essential for ensuring the reliability of downstream applications. However, existing automated data unit t…