3 citations · 4 across the 2 of their papers we have counts for
Showing cs.DBShow all
3 papers · 1 filter
cs.DB2024
SPADE: Synthesizing Data Quality Assertions for Large Language Model Pipelines
Shreya Shankar, Haotian Li, Parth Asawa +7
Large language models (LLMs) are being increasingly deployed as part of pipelines that repeatedly process or generate data of some sort. However, a common barrier to deployment are…
cs.DB2016★ 3 cited
Data Source Selection for Information Integration in Big Data Era
Yiming Lin, Hongzhi Wang, Jianzhong Li +1
In Big data era, information integration often requires abundant data extracted from massive data sources. Due to a large number of data sources, data source selection plays a cruc…
cs.DB2016★ 1 cited
Efficient Entity Resolution on Heterogeneous Records
Yiming Lin, Hongzhi Wang, Jianzhong Li +1
Entity resolution (ER) is the problem of identifying and merging records that refer to the same real-world entity. In many scenarios, raw records are stored under heterogeneous env…