collaborators

5 papers

cs.CL2026

Alignment-Guided Largest Table Overlap Size Estimation

Ge Lee, Shixun Huang, Zhifeng Bao +2

Fast estimation of the size of the largest overlap between tables enables blocking and query-by-table retrieval in large table repositories. The first and the state-of-the-art esti…

cs.DB2026

LEARNT: A Practical Estimator for Cardinality of LIKE Queries with Formal Accuracy Guarantees

Hai Lan, Zhifeng Bao, Divesh Srivastava +3

We study the problem of cardinality estimation for LIKE queries on string data, focusing on the most common patterns in real workloads: prefix, suffix, and substring queries. We pr…

cs.DB2026

Unified Data Discovery across Query Modalities and User Intents

Tingting Wang, Shixun Huang, Zhifeng Bao +4

Data discovery - retrieving relevant tables from a data lake in response to user queries - is a fundamental building block for downstream analytics. In practice, data discovery mus…

cs.DB2026

Shape-Agnostic Table Overlap Discovery: A Maximum Common Subhypergraph Approach

Ge Lee, Shixun Huang, Zhifeng Bao +3

Understanding how two tables overlap is useful for many data management tasks, but challenging because tables often differ in row and column orders and lack reliable metadata in pr…

cs.DB2025

Distinctiveness Maximization in Datasets Assemblage

Tingting Wang, Shixun Huang, Zhifeng Bao +3

In this paper, given a user's query set and budget, we aim to use the limited budget to help users assemble a set of datasets that can enrich a base dataset by introducing the maxi…