11 citations · 13 across the 3 of their papers we have counts for
3 papers · 1 filter
MergeDistill: Merging Pre-trained Language Models using Distillation
Simran Khanuja, Melvin Johnson, Partha Talukdar
Pre-trained multilingual language models (LMs) have achieved state-of-the-art results in cross-lingual transfer, but they often lead to an inequitable representation of languages d…
nmT5 -- Is parallel data still relevant for pre-training massively multilingual language models?
Mihir Kale, Aditya Siddhant, Noah Constant +3
Recently, mT5 - a massively multilingual version of T5 - leveraged a unified text-to-text format to attain state-of-the-art results on a wide variety of multilingual NLP tasks. In…
Gradient-guided Loss Masking for Neural Machine Translation
Xinyi Wang, Ankur Bapna, Melvin Johnson +1
To mitigate the negative effect of low quality training data on the performance of neural machine translation models, most existing strategies focus on filtering out harmful data b…