activity
20152022
most citedCodeBLEU: a Method for Automatic Evaluation of Code Synthesis

189 citations · 600 across the 18 of their papers we have counts for

collaborators
Showing cs.CLShow all

41 papers · 1 filter

cs.CL2020

Neural Deepfake Detection with Factual Structure of Text

Wanjun Zhong, Duyu Tang, Zenan Xu +5

Deepfake detection, the task of automatically discriminating machine-generated text, is increasingly critical with recent advances in natural language generative models. Existing a…

cs.CL2020

Improving the Efficiency of Grammatical Error Correction with Erroneous Span Detection and Correction

Mengyun Chen, Tao Ge, Xingxing Zhang +2

We propose a novel language-independent approach to improve the efficiency for Grammatical Error Correction (GEC) by dividing the task into two subtasks: Erroneous Span Detection (…

cs.CL20201 cited

Tell Me How to Ask Again: Question Data Augmentation with Controllable Rewriting in Continuous Space

Dayiheng Liu, Yeyun Gong, Jie Fu +5

In this paper, we propose a novel data augmentation method, referred to as Controllable Rewriting based Question Data Augmentation (CRQDA), for machine reading comprehension (MRC),…

cs.CL2020

InfoXLM: An Information-Theoretic Framework for Cross-Lingual Language Model Pre-Training

Zewen Chi, Li Dong, Furu Wei +7

In this work, we present an information-theoretic framework that formulates cross-lingual language model pre-training as maximizing mutual information between multilingual-multi-gr…

cs.CL2019

Improving Grammatical Error Correction with Machine Translation Pairs

Wangchunshu Zhou, Tao Ge, Chang Mu +3

We propose a novel data synthesis method to generate diverse error-corrected sentence pairs for improving grammatical error correction, which is based on a pair of machine translat…

cs.CL2019

Unicoder: A Universal Language Encoder by Pre-training with Multiple Cross-lingual Tasks

Haoyang Huang, Yaobo Liang, Nan Duan +4

We present Unicoder, a universal language encoder that is insensitive to different languages. Given an arbitrary NLP task, a model can be trained with Unicoder using training data…