6 papers · 1 filter
From Research Questions to Columns: Operationalization-Aware Data Discovery
Houming Chen, H. V. Jagadish
Researchers often approach a data repository with an abstract concept and ask which columns can measure it. Useful columns may not resemble the query; they may matter only as compl…
The General Stability of Ranking
Houming Chen, H. V. Jagadish
Rankings derived from weighted scoring functions are widely used in settings such as university rankings and employment candidate evaluations. Since ranking weights are often chose…
ConStruM: A Structure-Guided LLM Framework for Context-Aware Schema Matching
Houming Chen, Zhe Zhang, H. V. Jagadish
Column matching is a central task in reconciling schemas for data integration. Column names and descriptions are valuable for this task. LLMs can leverage such natural-language sch…
SemBench: A Benchmark for Semantic Query Processing Engines
Jiale Lao, Andreas Zimmerer, Olga Ovcharenko +12
We present a benchmark targeting a novel class of systems: semantic query processing engines. Those systems rely inherently on generative and reasoning capabilities of state-of-the…
Nexus: Inferring Join Graphs from Metadata Alone via Iterative Low-Rank Matrix Completion
Tianji Cong, Yuanyuan Tian, Andreas Mueller +5
Automatically inferring join relationships is a critical task for effective data discovery, integration, querying and reuse. However, accurately and efficiently identifying these r…
OpenForge: Probabilistic Metadata Integration
Tianji Cong, Fatemeh Nargesian, Junjie Xing +1
Modern data stores increasingly rely on metadata for enabling diverse activities such as data cataloging and search. However, metadata curation remains a labor-intensive task, and…