7 papers
Labeling Training Data for Entity Matching Using Large Language Models
Aaron Steiner, Christian Bizer
Recent large language models (LLMs) achieve strong performance on entity matching without requiring task-specific training data. However, applying these models to large sets of can…
Automatic End-to-End Data Integration using Large Language Models
Aaron Steiner, Christian Bizer
Designing data integration pipelines typically requires substantial manual effort from data engineers to configure pipeline components and label training data. While LLMs have show…
MCP vs RAG vs NLWeb vs HTML: A Comparison of the Effectiveness and Efficiency of Different Agent Interfaces to the Web (Technical Report)
Aaron Steiner, Ralph Peeters, Christian Bizer
Large language model agents are increasingly used to automate web tasks such as product search, offer comparison, and checkout. Current research explores different interfaces throu…
WebMall -- A Multi-Shop Benchmark for Evaluating Web Agents
Ralph Peeters, Aaron Steiner, Luca Schwarz +2
LLM-based web agents have the potential to automate long-running web tasks, such as searching for products in multiple e-shops and subsequently ordering the cheapest products that…
Evaluating Knowledge Generation and Self-Refinement Strategies for LLM-based Column Type Annotation
Keti Korini, Christian Bizer
Understanding the semantics of columns in relational tables is an important pre-processing step for indexing data lakes in order to provide rich data search. An approach to establi…
Self-Refinement Strategies for LLM-based Product Attribute Value Extraction
Alexander Brinkmann, Christian Bizer
Structured product data, in the form of attribute-value pairs, is essential for e-commerce platforms to support features such as faceted product search and attribute-based product…