63 citations · 77 across the 6 of their papers we have counts for
6 papers
Data Selection Curriculum for Neural Machine Translation
Tasnim Mohiuddin, Philipp Koehn, Vishrav Chaudhary +3
Neural Machine Translation (NMT) models are typically trained on heterogeneous data that are concatenated and randomly shuffled. However, not all of the training data are equally u…
Tricks for Training Sparse Translation Models
Dheeru Dua, Shruti Bhosale, Vedanuj Goswami +3
Multi-task learning with an unbalanced data distribution skews model learning towards high resource tasks, especially when model capacity is fixed and fully shared across all tasks…
Facebook AI WMT21 News Translation Task Submission
Chau Tran, Shruti Bhosale, James Cross +3
We describe Facebook's multilingual model submission to the WMT2021 shared task on news translation. We participate in 14 language directions: English to and from Czech, German, Ha…
BASE Layers: Simplifying Training of Large, Sparse Models
Mike Lewis, Shruti Bhosale, Tim Dettmers +2
We introduce a new balanced assignment of experts (BASE) layer for large language models that greatly simplifies existing high capacity sparse layers. Sparse layers can dramaticall…
Language Models not just for Pre-training: Fast Online Neural Noisy Channel Modeling
Shruti Bhosale, Kyra Yee, Sergey Edunov +1
Pre-training models on vast quantities of unlabeled data has emerged as an effective approach to improving accuracy on many NLP tasks. On the other hand, traditional machine transl…
Beyond English-Centric Multilingual Machine Translation
Angela Fan, Shruti Bhosale, Holger Schwenk +14
Existing work in translation demonstrated the potential of massively multilingual machine translation by training a single model able to translate between any pair of languages. Ho…