2 papers
cs.CL2022
Mask More and Mask Later: Efficient Pre-training of Masked Language Models by Disentangling the [MASK] Token
Baohao Liao, David Thulke, Sanjika Hewavitharana +2
The pre-training of masked language models (MLMs) consumes massive computation to achieve good results on downstream NLP tasks, resulting in a large carbon footprint. In the vanill…
cs.CL2018
Back-Translation Sampling by Targeting Difficult Words in Neural Machine Translation
Marzieh Fadaee, Christof Monz
Neural Machine Translation has achieved state-of-the-art performance for several language pairs using a combination of parallel and synthetic data. Synthetic data is often generate…