Modeling Coherence for Neural Machine Translation with Dynamic and Topic Caches
arXiv:1711.11221
Abstract
Sentences in a well-formed text are connected to each other via various links to form the cohesive structure of the text. Current neural machine translation (NMT) systems translate a text in a conventional sentence-by-sentence fashion, ignoring such cross-sentence links and dependencies. This may lead to generate an incoherent target text for a coherent source text. In order to handle this issue, we propose a cache-based approach to modeling coherence for neural machine translation by capturing contextual information either from recently translated sentences or the entire document. Particularly, we explore two types of caches: a dynamic cache, which stores words from the best translation hypotheses of preceding sentences, and a topic cache, which maintains a set of target-side topical words that are semantically related to the document to be translated. On this basis, we build a new layer to score target words in these two caches with a cache-based neural model. Here the estimated probabilities from the cache-based neural model are combined with NMT probabilities into the final word prediction probabilities via a gating mechanism. Finally, the proposed cache-based neural model is trained jointly with NMT system in an end-to-end manner. Experiments and analysis presented in this paper demonstrate that the proposed cache-based model achieves substantial improvements over several state-of-the-art SMT and NMT baselines.
Accepted by COLING2018,11 pages, 3 figures
References in corpus (1)
Cited by in corpus (21)
- A Comprehensive Survey of Grammar Error Correction
- Pretrained Language Models for Document-Level Neural Machine Translation
- A Survey on Document-level Neural Machine Translation: Methods and Evaluation
- Acquiring Knowledge from Pre-trained Model to Neural Machine Translation
- Neural Machine Translation: Challenges, Progress and Future
- Divide and Rule: Effective Pre-Training for Context-Aware Multi-Encoder Translation Models
- Towards Making the Most of Context in Neural Machine Translation
- A Multi-task Multi-stage Transitional Training Framework for Neural Chat Translation
- Modeling Discourse Structure for Document-level Neural Machine Translation
- Document-level Neural Machine Translation with Document Embeddings
- When and Why is Document-level Context Useful in Neural Machine Translation?
- Document Graph for Neural Machine Translation
- Document Sub-structure in Neural Machine Translation
- Dynamic Context Selection for Document-level Neural Machine Translation via Reinforcement Learning
- Context-Aware Monolingual Repair for Neural Machine Translation
- Document-level Neural Machine Translation with Associated Memory Network
- Scalable Cross Lingual Pivots to Model Pronoun Gender for Translation
- Towards Making the Most of Dialogue Characteristics for Neural Chat Translation
- Modeling Bilingual Conversational Characteristics for Neural Chat Translation
- Towards User-Driven Neural Machine Translation
- Lexically Cohesive Neural Machine Translation with Copy Mechanism