activity
20192022
most citedCCNet: Extracting High Quality Monolingual Datasets from Web Crawl Data

246 citations · 631 across the 16 of their papers we have counts for

collaborators

23 papers

cs.LG20226 cited

TorchScale: Transformers at Scale

Shuming Ma, Hongyu Wang, Shaohan Huang +8

Large Transformers have achieved state-of-the-art performance across many tasks. Most open-source libraries on scaling Transformers focus on improving training or inference with be…

cs.CL20223 cited

Beyond English-Centric Bitexts for Better Multilingual Language Representation Learning

Barun Patra, Saksham Singhal, Shaohan Huang +5

In this paper, we elaborate upon recipes for building multilingual representation models that are not only competitive with existing state-of-the-art models but are also more param…

cs.LG202213 cited

Foundation Transformers

Hongyu Wang, Shuming Ma, Shaohan Huang +12

A big convergence of model architectures across language, vision, speech, and multimodal is emerging. However, under the same name "Transformers", the above areas use different imp…

cs.CL20222 cited

Data Selection Curriculum for Neural Machine Translation

Tasnim Mohiuddin, Philipp Koehn, Vishrav Chaudhary +3

Neural Machine Translation (NMT) models are typically trained on heterogeneous data that are concatenated and randomly shuffled. However, not all of the training data are equally u…

cs.CL20221 cited

OCR Improves Machine Translation for Low-Resource Languages

Oana Ignat, Jean Maillard, Vishrav Chaudhary +1

We aim to investigate the performance of current OCR systems on low resource languages and low resource scripts. We introduce and make publicly available a novel benchmark, OCR4MT,…

cs.CL20212 cited

Alternative Input Signals Ease Transfer in Multilingual Machine Translation

Simeng Sun, Angela Fan, James Cross +4

Recent work in multilingual machine translation (MMT) has focused on the potential of positive transfer between languages, particularly cases where higher-resourced languages can b…