activity
20202025
most citedBASE Layers: Simplifying Training of Large, Sparse Models

63 citations · 77 across the 6 of their papers we have counts for

collaborators

6 papers

cs.CL20222 cited

Data Selection Curriculum for Neural Machine Translation

Tasnim Mohiuddin, Philipp Koehn, Vishrav Chaudhary +3

Neural Machine Translation (NMT) models are typically trained on heterogeneous data that are concatenated and randomly shuffled. However, not all of the training data are equally u…

cs.CL2021

Tricks for Training Sparse Translation Models

Dheeru Dua, Shruti Bhosale, Vedanuj Goswami +3

Multi-task learning with an unbalanced data distribution skews model learning towards high resource tasks, especially when model capacity is fixed and fully shared across all tasks…

cs.CL20219 cited

Facebook AI WMT21 News Translation Task Submission

Chau Tran, Shruti Bhosale, James Cross +3

We describe Facebook's multilingual model submission to the WMT2021 shared task on news translation. We participate in 14 language directions: English to and from Czech, German, Ha…

cs.CL202163 cited

BASE Layers: Simplifying Training of Large, Sparse Models

Mike Lewis, Shruti Bhosale, Tim Dettmers +2

We introduce a new balanced assignment of experts (BASE) layer for large language models that greatly simplifies existing high capacity sparse layers. Sparse layers can dramaticall…

cs.CL20202 cited

Language Models not just for Pre-training: Fast Online Neural Noisy Channel Modeling

Shruti Bhosale, Kyra Yee, Sergey Edunov +1

Pre-training models on vast quantities of unlabeled data has emerged as an effective approach to improving accuracy on many NLP tasks. On the other hand, traditional machine transl…

cs.CL2020

Beyond English-Centric Multilingual Machine Translation

Angela Fan, Shruti Bhosale, Holger Schwenk +14

Existing work in translation demonstrated the potential of massively multilingual machine translation by training a single model able to translate between any pair of languages. Ho…