55 citations · 73 across the 4 of their papers we have counts for
7 papers
Demonstration of Panda: A Weakly Supervised Entity Matching System
Renzhi Wu, Prem Sakala, Peng Li +2
Entity matching (EM) refers to the problem of identifying tuple pairs in one or more relations that refer to the same real world entities. Supervised machine learning (ML) approach…
Auto-FuzzyJoin: Auto-Program Fuzzy Similarity Joins Without Labeled Examples
Peng Li, Xiang Cheng, Xu Chu +2
Fuzzy similarity join is an important database operator widely used in practice. So far the research community has focused exclusively on optimizing fuzzy join \textit{scalability}…
Nearest Neighbor Classifiers over Incomplete Information: From Certain Answers to Certain Predictions
Bojan Karlaš, Peng Li, Renzhi Wu +4
Machine learning (ML) applications have been thriving recently, largely attributed to the increasing availability of data. However, inconsistency and incomplete information are ubi…
ZeroER: Entity Resolution using Zero Labeled Examples
Renzhi Wu, Sanya Chaba, Saurabh Sawlani +2
Entity resolution (ER) refers to the problem of matching records in one or more relations that refer to the same real-world entity. While supervised machine learning (ML) approache…
CleanML: A Study for Evaluating the Impact of Data Cleaning on ML Classification Tasks
Peng Li, Xi Rao, Jennifer Blase +3
Data quality affects machine learning (ML) model performances, and data scientists spend considerable amount of time on data cleaning before model training. However, to date, there…
GOGGLES: Automatic Image Labeling with Affinity Coding
Nilaksh Das, Sanya Chaba, Renzhi Wu +3
Generating large labeled training data is becoming the biggest bottleneck in building and deploying supervised machine learning models. Recently, the data programming paradigm has…