54 citations · 177 across the 15 of their papers we have counts for
20 papers · 1 filter
Beyond English-Centric Bitexts for Better Multilingual Language Representation Learning
Barun Patra, Saksham Singhal, Shaohan Huang +5
In this paper, we elaborate upon recipes for building multilingual representation models that are not only competitive with existing state-of-the-art models but are also more param…
CROP: Zero-shot Cross-lingual Named Entity Recognition with Multilingual Labeled Sequence Translation
Jian Yang, Shaohan Huang, Shuming Ma +6
Named entity recognition (NER) suffers from the scarcity of annotated training data, especially for low-resource languages without labeled data. Cross-lingual NER has been proposed…
DeepNet: Scaling Transformers to 1,000 Layers
Hongyu Wang, Shuming Ma, Li Dong +3
In this paper, we propose a simple yet effective method to stabilize extremely deep Transformers. Specifically, we introduce a new normalization function (DeepNorm) to modify the r…
Multilingual Machine Translation Systems from Microsoft for WMT21 Shared Task
Jian Yang, Shuming Ma, Haoyang Huang +8
This report describes Microsoft's machine translation systems for the WMT21 shared task on large-scale multilingual machine translation. We participated in all three evaluation tra…
Improving Non-autoregressive Generation with Mixup Training
Ting Jiang, Shaohan Huang, Zihan Zhang +6
While pre-trained language models have achieved great success on various natural language understanding tasks, how to effectively leverage them into non-autoregressive generation t…
Allocating Large Vocabulary Capacity for Cross-lingual Language Model Pre-training
Bo Zheng, Li Dong, Shaohan Huang +5
Compared to monolingual models, cross-lingual models usually require a more expressive vocabulary to represent all languages adequately. We find that many languages are under-repre…