activity
20192022
most citedSurfCon: Synonym Discovery on Privacy-Aware Clinical Data

10 citations · 14 across the 4 of their papers we have counts for

collaborators

8 papers

cs.CL2022

Synthetic Question Value Estimation for Domain Adaptation of Question Answering

Xiang Yue, Ziyu Yao, Huan Sun

Synthesizing QA pairs with a question generator (QG) on the target domain has become a popular approach for domain adaptation of question answering (QA) models. Since synthetic que…

cs.CL20212 cited

Differential Privacy for Text Analytics via Natural Text Sanitization

Xiang Yue, Minxin Du, Tianhao Wang +3

Texts convey sophisticated knowledge. However, texts also convey sensitive information. Despite the success of general-purpose language models and domain-specific mechanisms with d…

cs.CL2020

PHICON: Improving Generalization of Clinical Text De-identification Models via Data Augmentation

Xiang Yue, Shuang Zhou

De-identification is the task of identifying protected health information (PHI) in the clinical text. Existing neural de-identification models often fail to generalize to a new dat…

cs.CL2020

COUGH: A Challenge Dataset and Models for COVID-19 FAQ Retrieval

Xinliang Frederick Zhang, Heming Sun, Xiang Yue +2

We present a large, challenging dataset, COUGH, for COVID-19 FAQ retrieval. Similar to a standard FAQ dataset, COUGH consists of three parts: FAQ Bank, Query Bank and Relevance Set…

cs.CL20202 cited

Clinical Reading Comprehension: A Thorough Analysis of the emrQA Dataset

Xiang Yue, Bernal Jimenez Gutierrez, Huan Sun

Machine reading comprehension has made great progress in recent years owing to large-scale annotated datasets. In the clinical domain, however, creating such datasets is quite diff…

cs.CL2020

Practical Annotation Strategies for Question Answering Datasets

Bernhard Kratzwald, Xiang Yue, Huan Sun +1

Annotating datasets for question answering (QA) tasks is very costly, as it requires intensive manual labor and often domain-specific knowledge. Yet strategies for annotating QA da…