1 citations · 1 across the 5 of their papers we have counts for
5 papers
Automatic String Data Validation with Pattern Discovery
Xinwei Lin, Jing Zhao, Peng Di +7
In enterprise data pipelines, data insertions occur periodically and may impact downstream services if data quality issues are not addressed. Typically, such problems can be invest…
Antonym vs Synonym Distinction using InterlaCed Encoder NETworks (ICE-NET)
Muhammad Asif Ali, Yan Hu, Jianbin Qin +1
Antonyms vs synonyms distinction is a core challenge in lexico-semantic analysis and automated lexical resource construction. These pairs share a similar distributional context whi…
BClean: A Bayesian Data Cleaning System
Jianbin Qin, Sifan Huang, Yaoshu Wang +6
There is a considerable body of work on data cleaning which employs various principles to rectify erroneous data and transform a dirty dataset into a cleaner one. One of prevalent…
GARI: Graph Attention for Relative Isomorphism of Arabic Word Embeddings
Muhammad Asif Ali, Maha Alshmrani, Jianbin Qin +2
Bilingual Lexical Induction (BLI) is a core challenge in NLP, it relies on the relative isomorphism of individual embedding spaces. Existing attempts aimed at controlling the relat…
GRI: Graph-based Relative Isomorphism of Word Embedding Spaces
Muhammad Asif Ali, Yan Hu, Jianbin Qin +1
Automated construction of bilingual dictionaries using monolingual embedding spaces is a core challenge in machine translation. The end performance of these dictionaries relies upo…