Improving Neural Machine Translation with Pre-trained Representation
arXiv:1908.07688
Abstract
Monolingual data has been demonstrated to be helpful in improving the translation quality of neural machine translation (NMT). The current methods stay at the usage of word-level knowledge, such as generating synthetic parallel data or extracting information from word embedding. In contrast, the power of sentence-level contextual knowledge which is more complex and diverse, playing an important role in natural language generation, has not been fully exploited. In this paper, we propose a novel structure which could leverage monolingual data to acquire sentence-level contextual representations. Then, we design a framework for integrating both source and target sentence-level representations into NMT model to improve the translation quality. Experimental results on Chinese-English, German-English machine translation tasks show that our proposed model achieves improvement over strong Transformer baselines, while experiments on English-Turkish further demonstrate the effectiveness of our approach in the low-resource scenario.
In Progress
References in corpus (10)
- Sequence to Sequence Learning with Neural Networks
- On Using Monolingual Corpora in Neural Machine Translation
- Distilling Task-Specific Knowledge from BERT into Simple Neural Networks
- Investigating Backtranslation in Neural Machine Translation
- DiSAN: Directional Self-Attention Network for RNN/CNN-Free Language Understanding
- On the State of the Art of Evaluation in Neural Language Models
- Multi-layer Representation Fusion for Neural Machine Translation
- Why Self-Attention? A Targeted Evaluation of Neural Machine Translation Architectures
- Exploiting Deep Representations for Neural Machine Translation
- Beyond Error Propagation in Neural Machine Translation: Characteristics of Language Also Matter