activity
20172022
most citedAuto-FuzzyJoin: Auto-Program Fuzzy Similarity Joins Without Labeled Examples

7 citations · 15 across the 6 of their papers we have counts for

collaborators

6 papers

cs.DB20221 cited

Ground Truth Inference for Weakly Supervised Entity Matching

Renzhi Wu, Alexander Bendeck, Xu Chu +1

Entity matching (EM) refers to the problem of identifying pairs of data records in one or more relational tables that refer to the same entity in the real world. Supervised machine…

cs.DB20211 cited

Auto-Pipeline: Synthesizing Complex Data Pipelines By-Target Using Reinforcement Learning and Search

Junwen Yang, Yeye He, Surajit Chaudhuri

Recent work has made significant progress in helping users to automate single data preparation steps, such as string-transformations and table-manipulation operators (e.g., Join, G…

cs.DB20214 cited

Demonstration of Panda: A Weakly Supervised Entity Matching System

Renzhi Wu, Prem Sakala, Peng Li +2

Entity matching (EM) refers to the problem of identifying tuple pairs in one or more relations that refer to the same real world entities. Supervised machine learning (ML) approach…

cs.DB20212 cited

Auto-Validate: Unsupervised Data Validation Using Data-Domain Patterns Inferred from Data Lakes

Jie Song, Yeye He

Complex data pipelines are increasingly common in diverse applications such as BI reporting and ML modeling. These pipelines often recur regularly (e.g., daily or weekly), as BI re…

cs.DB20217 cited

Auto-FuzzyJoin: Auto-Program Fuzzy Similarity Joins Without Labeled Examples

Peng Li, Xiang Cheng, Xu Chu +2

Fuzzy similarity join is an important database operator widely used in practice. So far the research community has focused exclusively on optimizing fuzzy join \textit{scalability}…

cs.DB2017

Synthesizing Mapping Relationships Using Table Corpus

Yue Wang, Yeye He

Mapping relationships, such as (country, country-code) or (company, stock-ticker), are versatile data assets for an array of applications in data cleaning and data integration like…