collaborators

5 papers

cs.DB2026

MosaicJoin: Compact Semantic Sketches for Value-Level Join Discovery

Grace Fan, Eden Wu, Majid Daliri +1

Join discovery is a core task in dataset search, enabling users to find columns that can be joined with a given query column. Early approaches focused on equi-joins, but data lakes…

cs.CL2026

LakeQA: An Exploratory QA Benchmark over a Million-Scale Data Lake

Haonan Wang, Jiaxiang Liu, Yurong Liu +11

Recent large language models (LLMs) have shown rapid progress in reading-based question answering (QA), where evidence is explicitly provided or can be trivially retrieved. In cont…

cs.DB2025

AutoDDG: Automated Dataset Description Generation using Large Language Models

Haoxiang Zhang, Yurong Liu, Aécio Santos +2

The proliferation of datasets across open data portals and enterprise data lakes presents an opportunity for deriving data-driven insights. Widely-used dataset search systems rely…

cs.AI2025

Interactive Data Harmonization with LLM Agents: Opportunities and Challenges

Aécio Santos, Eduardo H. M. Pena, Roque Lopez +1

Data harmonization is an essential task that entails integrating datasets from diverse sources. Despite years of research in this area, it remains a time-consuming and challenging…

cs.DB2025

Magneto: Combining Small and Large Language Models for Schema Matching

Yurong Liu, Eduardo Pena, Aecio Santos +2

Recent advances in language models opened new opportunities to address complex schema matching tasks. Schema matching approaches have been proposed that demonstrate the usefulness…