activity
20122021
most citedCrowdER: Crowdsourcing Entity Resolution

66 citations · 138 across the 7 of their papers we have counts for

collaborators

10 papers

cs.LG20212 cited

Enabling SQL-based Training Data Debugging for Federated Learning

Yejia Liu, Weiyuan Wu, Lampros Flokas +2

How can we debug a logistical regression model in a federated learning setting when seeing the model behave unexpectedly (e.g., the model rejects all high-income customers' loan ap…

cs.DB202129 cited

DataPrep.EDA: Task-Centric Exploratory Data Analysis for Statistical Modeling in Python

Jinglin Peng, Weiyuan Wu, Brandon Lockhart +6

Exploratory Data Analysis (EDA) is a crucial step in any data science project. However, existing Python libraries fall short in supporting data scientists to complete common EDA ta…

cs.DB2021

Explaining Inference Queries with Bayesian Optimization

Brandon Lockhart, Jinglin Peng, Weiyuan Wu +2

Obtaining an explanation for an SQL query result can enrich the analysis experience, reveal data errors, and provide deeper insight into the data. Inference query explanation seeks…

cs.DB2020

Are We Ready For Learned Cardinality Estimation?

Xiaoying Wang, Changbo Qu, Weiyuan Wu +2

Cardinality estimation is a fundamental but long unresolved problem in query optimization. Recently, multiple papers from different research groups consistently report that learned…

cs.DB202036 cited

Complaint-driven Training Data Debugging for Query 2.0

Weiyuan Wu, Lampros Flokas, Eugene Wu +1

As the need for machine learning (ML) increases rapidly across all industry sectors, there is a significant interest among commercial database providers to support "Query 2.0", whi…

cs.DB20192 cited

Towards Extracting Highlights From Recorded Live Videos: An Implicit Crowdsourcing Approach

Ruochen Jiang, Changbo Qu, Jiannan Wang +2

Live streaming platforms need to store a lot of recorded live videos on a daily basis. An important problem is how to automatically extract highlights (i.e., attractive short video…