Neural Machine Translation with Monolingual Translation Memory
arXiv:2105.11269
Abstract
Prior work has proved that Translation memory (TM) can boost the performance of Neural Machine Translation (NMT). In contrast to existing work that uses bilingual corpus as TM and employs source-side similarity search for memory retrieval, we propose a new framework that uses monolingual memory and performs learnable memory retrieval in a cross-lingual manner. Our framework has unique advantages. First, the cross-lingual memory retriever allows abundant monolingual data to be TM. Second, the memory retriever and NMT model can be jointly optimized for the ultimate translation goal. Experiments show that the proposed method obtains substantial improvements. Remarkably, it even outperforms strong TM-augmented NMT baselines using bilingual TM. Owning to the ability to leverage monolingual data, our model also demonstrates effectiveness in low-resource and domain adaptation scenarios.
ACL2021
References in corpus (9)
- Language Models are Few-Shot Learners
- Dual Learning for Machine Translation
- REALM: Retrieval-Augmented Language Model Pre-Training
- On Using Monolingual Corpora in Neural Machine Translation
- Asymmetric LSH (ALSH) for Sublinear Time Maximum Inner Product Search (MIPS)
- Pre-training via Paraphrasing
- TranSmart: A Practical Interactive Machine Translation System
- Learning to Reuse Translations: Guiding Neural Machine Translation with Examples
- Self-Training Sampling with Monolingual Data Uncertainty for Neural Machine Translation