7 citations · 15 across the 6 of their papers we have counts for
6 papers
Ground Truth Inference for Weakly Supervised Entity Matching
Renzhi Wu, Alexander Bendeck, Xu Chu +1
Entity matching (EM) refers to the problem of identifying pairs of data records in one or more relational tables that refer to the same entity in the real world. Supervised machine…
Auto-Pipeline: Synthesizing Complex Data Pipelines By-Target Using Reinforcement Learning and Search
Junwen Yang, Yeye He, Surajit Chaudhuri
Recent work has made significant progress in helping users to automate single data preparation steps, such as string-transformations and table-manipulation operators (e.g., Join, G…
Demonstration of Panda: A Weakly Supervised Entity Matching System
Renzhi Wu, Prem Sakala, Peng Li +2
Entity matching (EM) refers to the problem of identifying tuple pairs in one or more relations that refer to the same real world entities. Supervised machine learning (ML) approach…
Auto-Validate: Unsupervised Data Validation Using Data-Domain Patterns Inferred from Data Lakes
Jie Song, Yeye He
Complex data pipelines are increasingly common in diverse applications such as BI reporting and ML modeling. These pipelines often recur regularly (e.g., daily or weekly), as BI re…
Auto-FuzzyJoin: Auto-Program Fuzzy Similarity Joins Without Labeled Examples
Peng Li, Xiang Cheng, Xu Chu +2
Fuzzy similarity join is an important database operator widely used in practice. So far the research community has focused exclusively on optimizing fuzzy join \textit{scalability}…
Synthesizing Mapping Relationships Using Table Corpus
Yue Wang, Yeye He
Mapping relationships, such as (country, country-code) or (company, stock-ticker), are versatile data assets for an array of applications in data cleaning and data integration like…