activity
20172023
most citedCodeXGLUE: A Machine Learning Benchmark Dataset for Code Understanding and Generation

416 citations · 855 across the 28 of their papers we have counts for

collaborators
Showing 2022Show all

9 papers · 1 filter

cs.CL2022★ 9 cited

Effidit: Your AI Writing Assistant

Shuming Shi, Enbo Zhao, Duyu Tang +10

In this technical report, we introduce Effidit (Efficient and Intelligent Editing), a digital writing assistant that facilitates users to write higher-quality text more efficiently…

cs.CL2022★ 7 cited

One Model, Multiple Modalities: A Sparsely Activated Approach for Text, Sound, Image, Video and Code

Yong Dai, Duyu Tang, Liangxin Liu +7

People perceive the world with multiple senses (e.g., through hearing sounds, reading words and seeing objects). However, most existing AI systems only process an individual modali…

cs.CL2022

SkillNet-NLG: General-Purpose Natural Language Generation with a Sparsely Activated Approach

Junwei Liao, Duyu Tang, Fan Zhang +1

We present SkillNet-NLG, a sparsely activated approach that handles many natural language generation tasks with one model. Different from traditional dense models that always activ…

cs.CL2022★ 1 cited

Pretraining Chinese BERT for Detecting Word Insertion and Deletion Errors

Cong Zhou, Yong Dai, Duyu Tang +4

Chinese BERT models achieve remarkable progress in dealing with grammatical errors of word substitution. However, they fail to handle word insertion and deletion because BERT assum…

cs.CL2022★ 1 cited

"Is Whole Word Masking Always Better for Chinese BERT?": Probing on Chinese Grammatical Error Correction

Yong Dai, Linyang Li, Cong Zhou +5

Whole word masking (WWM), which masks all subwords corresponding to a word at once, makes a better English BERT model. For the Chinese language, however, there is no subword becaus…

cs.CL2022

Exploring and Adapting Chinese GPT to Pinyin Input Method

Minghuan Tan, Yong Dai, Duyu Tang +5

While GPT has become the de-facto method for text generation tasks, its application to pinyin input method remains unexplored. In this work, we make the first exploration to levera…