7 papers
From Research Questions to Columns: Operationalization-Aware Data Discovery
Houming Chen, H. V. Jagadish
Researchers often approach a data repository with an abstract concept and ask which columns can measure it. Useful columns may not resemble the query; they may matter only as compl…
The General Stability of Ranking
Houming Chen, H. V. Jagadish
Rankings derived from weighted scoring functions are widely used in settings such as university rankings and employment candidate evaluations. Since ranking weights are often chose…
ConStruM: A Structure-Guided LLM Framework for Context-Aware Schema Matching
Houming Chen, Zhe Zhang, H. V. Jagadish
Column matching is a central task in reconciling schemas for data integration. Column names and descriptions are valuable for this task. LLMs can leverage such natural-language sch…
SemBench: A Benchmark for Semantic Query Processing Engines
Jiale Lao, Andreas Zimmerer, Olga Ovcharenko +12
We present a benchmark targeting a novel class of systems: semantic query processing engines. Those systems rely inherently on generative and reasoning capabilities of state-of-the…
MMTU: A Massive Multi-Task Table Understanding and Reasoning Benchmark
Junjie Xing, Yeye He, Mengyu Zhou +6
Tables and table-based use cases play a crucial role in many important real-world applications, such as spreadsheets, databases, and computational notebooks, which traditionally re…
Nexus: Inferring Join Graphs from Metadata Alone via Iterative Low-Rank Matrix Completion
Tianji Cong, Yuanyuan Tian, Andreas Mueller +5
Automatically inferring join relationships is a critical task for effective data discovery, integration, querying and reuse. However, accurately and efficiently identifying these r…