2 papers
cs.DB2023
Auto-Validate by-History: Auto-Program Data Quality Constraints to Validate Recurring Data Pipelines
Dezhan Tu, Yeye He, Weiwei Cui +5
Data pipelines are widely employed in modern enterprises to power a variety of Machine-Learning (ML) and Business-Intelligence (BI) applications. Crucially, these pipelines are \em…
cs.DB2021
Auto-Tag: Tagging-Data-By-Example in Data Lakes
Yeye He, Jie Song, Yue Wang +5
As data lakes become increasingly popular in large enterprises today, there is a growing need to tag or classify data assets (e.g., files and databases) in data lakes with addition…