29 citations · 30 across the 3 of their papers we have counts for
4 papers
A Survey on Data Cleaning Methods for Improved Machine Learning Model Performance
Ga Young Lee, Lubna Alzamil, Bakhtiyar Doskenov +1
Data cleaning is the initial stage of any machine learning project and is one of the most critical processes in data analysis. It is a critical step in ensuring that the dataset is…
Learning Over Dirty Data Without Cleaning
Jose Picado, John Davis, Arash Termehchy +1
Real-world datasets are dirty and contain many errors. Examples of these issues are violations of integrity constraints, duplicates, and inconsistencies in representing data values…
Managing Variability in Relational Databases by VDBMS
Parisa Ataei, Qiaoran Li, Eric Walkingshaw +1
Variability inherently exists in databases in various contexts which creates database variants. For example, variants of a database could have different schemas/content (database e…
Integrating Information About Entities Progressively
Ben McCamish, Christopher Buss, Arash Termehchy +1
Users often have to integrate information about entities from multiple data sources. This task is challenging as each data source may represent information about the same entity in…