10 citations · 14 across the 4 of their papers we have counts for
8 papers
Synthetic Question Value Estimation for Domain Adaptation of Question Answering
Xiang Yue, Ziyu Yao, Huan Sun
Synthesizing QA pairs with a question generator (QG) on the target domain has become a popular approach for domain adaptation of question answering (QA) models. Since synthetic que…
Differential Privacy for Text Analytics via Natural Text Sanitization
Xiang Yue, Minxin Du, Tianhao Wang +3
Texts convey sophisticated knowledge. However, texts also convey sensitive information. Despite the success of general-purpose language models and domain-specific mechanisms with d…
PHICON: Improving Generalization of Clinical Text De-identification Models via Data Augmentation
Xiang Yue, Shuang Zhou
De-identification is the task of identifying protected health information (PHI) in the clinical text. Existing neural de-identification models often fail to generalize to a new dat…
COUGH: A Challenge Dataset and Models for COVID-19 FAQ Retrieval
Xinliang Frederick Zhang, Heming Sun, Xiang Yue +2
We present a large, challenging dataset, COUGH, for COVID-19 FAQ retrieval. Similar to a standard FAQ dataset, COUGH consists of three parts: FAQ Bank, Query Bank and Relevance Set…
Clinical Reading Comprehension: A Thorough Analysis of the emrQA Dataset
Xiang Yue, Bernal Jimenez Gutierrez, Huan Sun
Machine reading comprehension has made great progress in recent years owing to large-scale annotated datasets. In the clinical domain, however, creating such datasets is quite diff…
Practical Annotation Strategies for Question Answering Datasets
Bernhard Kratzwald, Xiang Yue, Huan Sun +1
Annotating datasets for question answering (QA) tasks is very costly, as it requires intensive manual labor and often domain-specific knowledge. Yet strategies for annotating QA da…