49 citations · 140 across the 8 of their papers we have counts for
7 papers · 1 filter
Toward Data Cleaning with a Target Accuracy: A Case Study for Value Normalization
Adel Ardalan, Derek Paulsen, Amanpreet Singh Saini +2
Many applications need to clean data with a target accuracy. As far as we know, this problem has not been studied in depth. In this paper we take the first step toward solving it.…
Deep Entity Matching with Pre-Trained Language Models
Yuliang Li, Jinfeng Li, Yoshihiko Suhara +2
We present Ditto, a novel entity matching system based on pre-trained Transformer-based language models. We fine-tune and cast EM as a sequence-pair classification problem to lever…
Data Curation with Deep Learning [Vision]
Saravanan Thirumuruganathan, Nan Tang, Mourad Ouzzani +1
Data curation - the process of discovering, integrating, and cleaning data - is one of the oldest, hardest, yet inevitable data management problems. Despite decades of efforts from…
Toward a System Building Agenda for Data Integration
AnHai Doan, Adel Ardalan, Jeffrey R. Ballard +7
In this paper we argue that the data management community should devote far more effort to building data integration (DI) systems, in order to truly advance the field. Toward this…
Muppet: MapReduce-Style Processing of Fast Data
Wang Lam, Lu Liu, STS Prasad +3
MapReduce has emerged as a popular method to process big data. In the past few years, however, not just big data, but fast data has also exploded in volume and availability. Exampl…
Tuffy: Scaling up Statistical Inference in Markov Logic Networks using an RDBMS
Feng Niu, Christopher Ré, AnHai Doan +1
Markov Logic Networks (MLNs) have emerged as a powerful framework that combines statistical and logical reasoning; they have been applied to many data intensive problems including…