6 papers
MaDI-Bench: An End-to-End Data Integration Benchmark
Aaron Steiner, Ralph Peeters, Christian Bizer
Data integration combines heterogeneous data sets into a single, coherent representation. Data integration involves a sequence of interdependent tasks including schema matching, va…
Labeling Training Data for Entity Matching Using Large Language Models
Aaron Steiner, Christian Bizer
Recent large language models (LLMs) achieve strong performance on entity matching without requiring task-specific training data. However, applying these models to large sets of can…
WebMall -- A Multi-Shop Benchmark for Evaluating Web Agents
Ralph Peeters, Aaron Steiner, Luca Schwarz +2
LLM-based web agents have the potential to automate long-running web tasks, such as searching for products in multiple e-shops and subsequently ordering the cheapest products that…
Automatic End-to-End Data Integration using Large Language Models
Aaron Steiner, Christian Bizer
Designing data integration pipelines typically requires substantial manual effort from data engineers to configure pipeline components and label training data. While LLMs have show…
MCP vs RAG vs NLWeb vs HTML: A Comparison of the Effectiveness and Efficiency of Different Agent Interfaces to the Web (Technical Report)
Aaron Steiner, Ralph Peeters, Christian Bizer
Large language model agents are increasingly used to automate web tasks such as product search, offer comparison, and checkout. Current research explores different interfaces throu…
Fine-tuning Large Language Models for Entity Matching
Aaron Steiner, Ralph Peeters, Christian Bizer
Generative large language models (LLMs) are a promising alternative to pre-trained language models for entity matching due to their high zero-shot performance and ability to genera…