activity
20192021
most citedXLM-T: Scaling up Multilingual Machine Translation with Pretrained Cross-lingual Transformer Encoders

23 citations · 38 across the 4 of their papers we have counts for

collaborators

7 papers

cs.CL202110 cited

A Robust and Domain-Adaptive Approach for Low-Resource Named Entity Recognition

Houjin Yu, Xian-Ling Mao, Zewen Chi +2

Recently, it has attracted much attention to build reliable named entity recognition (NER) systems using limited annotated data. Nearly all existing works heavily rely on domain-sp…

cs.CL202023 cited

XLM-T: Scaling up Multilingual Machine Translation with Pretrained Cross-lingual Transformer Encoders

Shuming Ma, Jian Yang, Haoyang Huang +10

Multilingual machine translation enables a single model to translate between different languages. Most existing multilingual machine translation systems adopt a randomly initialize…

cs.CL2020

Generating Informative Dialogue Responses with Keywords-Guided Networks

Heng-Da Xu, Xian-Ling Mao, Zewen Chi +3

Recently, open-domain dialogue systems have attracted growing attention. Most of them use the sequence-to-sequence (Seq2Seq) architecture to generate responses. However, traditiona…

cs.CL2020

InfoXLM: An Information-Theoretic Framework for Cross-Lingual Language Model Pre-Training

Zewen Chi, Li Dong, Furu Wei +7

In this work, we present an information-theoretic framework that formulates cross-lingual language model pre-training as maximizing mutual information between multilingual-multi-gr…

cs.CL20195 cited

Can Monolingual Pretrained Models Help Cross-Lingual Classification?

Zewen Chi, Li Dong, Furu Wei +2

Multilingual pretrained language models (such as multilingual BERT) have achieved impressive results for cross-lingual transfer. However, due to the constant model capacity, multil…

cs.CL2019

Cross-Lingual Natural Language Generation via Pre-Training

Zewen Chi, Li Dong, Furu Wei +3

In this work we focus on transferring supervision signals of natural language generation (NLG) tasks between multiple languages. We propose to pretrain the encoder and the decoder…