activity
20192025
most citedA Survey on Data Cleaning Methods for Improved Machine Learning Model Performance

29 citations · 30 across the 4 of their papers we have counts for

collaborators
Showing cs.DBShow all

7 papers · 1 filter

cs.DB2025

Towards Scalable Schema Mapping using Large Language Models

Christopher Buss, Mahdis Safari, Arash Termehchy +2

The growing need to integrate information from a large number of diverse sources poses significant scalability challenges for data integration systems. These systems often rely on…

cs.DB2023

Towards Consistent Language Models Using Declarative Constraints

Jasmin Mousavi, Arash Termehchy

Large language models have shown unprecedented abilities in generating linguistically coherent and syntactically correct natural language output. However, they often return incorre…

cs.DB2023

Multi-Agent Join

Vahid Ghadakchi, Mian Xie, Arash Termehchy +4

It is crucial to provide real-time performance in many applications, such as interactive and exploratory data analysis. In these settings, users often need to view subsets of query…

cs.DB202129 cited

A Survey on Data Cleaning Methods for Improved Machine Learning Model Performance

Ga Young Lee, Lubna Alzamil, Bakhtiyar Doskenov +1

Data cleaning is the initial stage of any machine learning project and is one of the most critical processes in data analysis. It is a critical step in ensuring that the dataset is…

cs.DB2020

Learning Over Dirty Data Without Cleaning

Jose Picado, John Davis, Arash Termehchy +1

Real-world datasets are dirty and contain many errors. Examples of these issues are violations of integrity constraints, duplicates, and inconsistencies in representing data values…

cs.DB20191 cited

Managing Variability in Relational Databases by VDBMS

Parisa Ataei, Qiaoran Li, Eric Walkingshaw +1

Variability inherently exists in databases in various contexts which creates database variants. For example, variants of a database could have different schemas/content (database e…