5 papers
Quantifying Retriever-Generator Alignment in RAG with Local Explanations
Korbinian Randl, Guido Rocchietti, Aron Henriksson +3
Retrieval-Augmented Generation (RAG) systems combine dense retrievers and language models to ground their outputs in external documents. However, the interaction between these comp…
Clean Me If You Can: A Large Collection of Real-World Addresses for Data Cleaning Benchmarking
Fatemeh Ahmadi, Tobias Bernhard, Mohamed Abdelmaksoud +3
There has been extensive research on automating and scaling data cleaning, i.e., the detection and correction of erroneous values in tabular data. Yet, existing approaches often pe…
RAMSeS: Robust and Adaptive Model Selection for Time-Series Anomaly Detection Algorithms
Mohamed Abdelmaksoud, Sheng Ding, Andrey Morozov +1
Time-series data vary widely across domains, making a universal anomaly detector impractical. Methods that perform well on one dataset often fail to transfer because what counts as…
Blend: A Unified Data Discovery System
Mahdi Esmailoghli, Christoph Schnell, Renée J. Miller +1
Most research on data discovery has so far focused on improving individual discovery operators such as join, correlation, or union discovery. However, in practice, a combination of…
Guiding Catalogue Enrichment with User Queries
Yupei Du, Jacek Golebiowski, Philipp Schmidt +1
Techniques for knowledge graph (KGs) enrichment have been increasingly crucial for commercial applications that rely on evolving product catalogues. However, because of the huge se…