activity
20172021
most citedHoloClean: Holistic Data Repairs with Probabilistic Inference

55 citations · 73 across the 4 of their papers we have counts for

collaborators

7 papers

cs.DB20214 cited

Demonstration of Panda: A Weakly Supervised Entity Matching System

Renzhi Wu, Prem Sakala, Peng Li +2

Entity matching (EM) refers to the problem of identifying tuple pairs in one or more relations that refer to the same real world entities. Supervised machine learning (ML) approach…

cs.DB20217 cited

Auto-FuzzyJoin: Auto-Program Fuzzy Similarity Joins Without Labeled Examples

Peng Li, Xiang Cheng, Xu Chu +2

Fuzzy similarity join is an important database operator widely used in practice. So far the research community has focused exclusively on optimizing fuzzy join \textit{scalability}…

cs.LG20207 cited

Nearest Neighbor Classifiers over Incomplete Information: From Certain Answers to Certain Predictions

Bojan Karlaš, Peng Li, Renzhi Wu +4

Machine learning (ML) applications have been thriving recently, largely attributed to the increasing availability of data. However, inconsistency and incomplete information are ubi…

cs.DB2019

ZeroER: Entity Resolution using Zero Labeled Examples

Renzhi Wu, Sanya Chaba, Saurabh Sawlani +2

Entity resolution (ER) refers to the problem of matching records in one or more relations that refer to the same real-world entity. While supervised machine learning (ML) approache…

cs.DB2019

CleanML: A Study for Evaluating the Impact of Data Cleaning on ML Classification Tasks

Peng Li, Xi Rao, Jennifer Blase +3

Data quality affects machine learning (ML) model performances, and data scientists spend considerable amount of time on data cleaning before model training. However, to date, there…

cs.CV2019

GOGGLES: Automatic Image Labeling with Affinity Coding

Nilaksh Das, Sanya Chaba, Renzhi Wu +3

Generating large labeled training data is becoming the biggest bottleneck in building and deploying supervised machine learning models. Recently, the data programming paradigm has…