Modeling Coverage for Neural Machine Translation
arXiv:1601.04811
Abstract
Attention mechanism has enhanced state-of-the-art Neural Machine Translation (NMT) by jointly learning to align and translate. It tends to ignore past alignment information, however, which often leads to over-translation and under-translation. To address this problem, we propose coverage-based NMT in this paper. We maintain a coverage vector to keep track of the attention history. The coverage vector is fed to the attention model to help adjust future attention, which lets NMT system to consider more about untranslated source words. Experiments show that the proposed approach significantly improves both translation quality and alignment quality over standard attention-based NMT.
Add subjective evaluation on top of ACL version: 25% of source words are under-translated by NMT
References in corpus (4)
- Sequence to Sequence Learning with Neural Networks
- Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling
- Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation
- Implicit Distortion and Fertility Models for Attention-based Encoder-Decoder NMT Model
Cited by in corpus (97)
- Achieving Human Parity on Automatic Chinese to English News Translation
- Language to Logical Form with Neural Attention
- Get To The Point: Summarization with Pointer-Generator Networks
- A Survey of Deep Learning Techniques for Neural Machine Translation
- RETURNN as a Generic Flexible Neural Toolkit with Application to Translation and Speech Recognition
- Syntactically Guided Neural Machine Translation
- Adversarial Neural Machine Translation
- Graph Neural Networks for Natural Language Processing: A Survey
- Massive Exploration of Neural Machine Translation Architectures
- Neural Machine Translation with Reconstruction
- Multilingual Extractive Reading Comprehension by Runtime Machine Translation
- Towards better decoding and language model integration in sequence to sequence models
- Neural Machine Translation with External Phrase Memory
- A Unified Query-based Generative Model for Question Generation and Question Answering
- Text Generation from Knowledge Graphs with Graph Transformers
- Conditional Variational Autoencoder for Neural Machine Translation
- Generative Neural Machine Translation
- Evaluating prose style transfer with the Bible
- Search Engine Guided Non-Parametric Neural Machine Translation
- Neural Text Generation: A Practical Guide
- Asynchronous Bidirectional Decoding for Neural Machine Translation
- Exploiting Cross-Sentence Context for Neural Machine Translation
- Neural Machine Translation with Pivot Languages
- Joint Training for Neural Machine Translation Models with Monolingual Data
- A Graph-to-Sequence Model for AMR-to-Text Generation
- Semi-Supervised QA with Generative Domain-Adaptive Nets
- Context-Aware Self-Attention Networks
- Neural Machine Translation with Supervised Attention
- Memory-enhanced Decoder for Neural Machine Translation
- Exploiting Deep Representations for Neural Machine Translation
- Modeling Source Syntax for Neural Machine Translation
- One Sentence One Model for Neural Machine Translation
- Multi-Head Attention with Disagreement Regularization
- Paraphrase Generation with Deep Reinforcement Learning
- Multi-Scale Attention with Dense Encoder for Handwritten Mathematical Expression Recognition
- Confidence through Attention
- Exploration on Generating Traditional Chinese Medicine Prescription from Symptoms with an End-to-End method
- Non-Autoregressive Machine Translation with Auxiliary Regularization
- Topic Augmented Generator for Abstractive Summarization
- Neural System Combination for Machine Translation
- Can Active Memory Replace Attention?
- Neural Machine Translation with Latent Semantic of Image and Text
- Context Gates for Neural Machine Translation
- Decoding-History-Based Adaptive Control of Attention for Neural Machine Translation
- Trainable Greedy Decoding for Neural Machine Translation
- Keyphrase Extraction with Span-based Feature Representations
- Natural Question Generation with Reinforcement Learning Based Graph-to-Sequence Model
- Learning to Remember Translation History with a Continuous Cache
- Source-side Prediction for Neural Headline Generation
- Prior Knowledge Integration for Neural Machine Translation using Posterior Regularization
- Sequential Interpretability: Methods, Applications, and Future Direction for Understanding Deep Learning Models in the Context of Sequential Data
- A GRU-based Encoder-Decoder Approach with Attention for Online Handwritten Mathematical Expression Recognition
- Variational Recurrent Neural Machine Translation
- AdaptSum: Towards Low-Resource Domain Adaptation for Abstractive Summarization
- Translating Phrases in Neural Machine Translation
- Improving Neural Machine Translation through Phrase-based Forced Decoding
- Segregated Temporal Assembly Recurrent Networks for Weakly Supervised Multiple Action Detection
- Neural Machine Translation Advised by Statistical Machine Translation
- Neural Machine Translation with Key-Value Memory-Augmented Attention
- Keyphrase Generation with Correlation Constraints
- Diverse Beam Search for Increased Novelty in Abstractive Summarization
- Recognizing Handwritten Mathematical Expressions as LaTex Sequences Using a Multiscale Robust Neural Network
- Data-to-Text Generation with Style Imitation
- LadRa-Net: Locally-Aware Dynamic Re-read Attention Net for Sentence Semantic Matching
- A study of latent monotonic attention variants
- Trajectory-based Radical Analysis Network for Online Handwritten Chinese Character Recognition
- Stroke Constrained Attention Network for Online Handwritten Mathematical Expression Recognition
- Multi-Domain Dialogue Acts and Response Co-Generation
- Energy-based Self-attentive Learning of Abstractive Communities for Spoken Language Understanding
- Testing Untestable Neural Machine Translation: An Industrial Case
- Chinese-Portuguese Machine Translation: A Study on Building Parallel Corpora from Comparable Texts
- Learning to Discriminate Noises for Incorporating External Information in Neural Machine Translation
- Radical analysis network for zero-shot learning in printed Chinese character recognition
- Neural Rule-Execution Tracking Machine For Transformer-Based Text Generation
- Learning Comment Generation by Leveraging User-Generated Data
- Modeling Future Cost for Neural Machine Translation
- Attention Focusing for Neural Machine Translation by Bridging Source and Target Embeddings
- Semantic Graphs for Generating Deep Questions
- PoinT-5: Pointer Network and T-5 based Financial NarrativeSummarisation
- English-Japanese Neural Machine Translation with Encoder-Decoder-Reconstructor
- Towards one-shot learning for rare-word translation with external experts
- Long Short-Term Memory with Dynamic Skip Connections
- Look-ahead Attention for Generation in Neural Machine Translation
- Neural Machine Translation: A Review and Survey
- Modeling Past and Future for Neural Machine Translation
- Neural Machine Translation Model with a Large Vocabulary Selected by Branching Entropy
- MeetSum: Transforming Meeting Transcript Summarization using Transformers!
- Word, Subword or Character? An Empirical Study of Granularity in Chinese-English NMT
- On Tree-Based Neural Sentence Modeling
- Future-Prediction-Based Model for Neural Machine Translation
- Learning Energy-Based Approximate Inference Networks for Structured Applications in NLP
- Learning to Generate Structured Queries from Natural Language with Indirect Supervision
- Abstractive Summarization Improved by WordNet-based Extractive Sentences
- Learning When to Concentrate or Divert Attention: Self-Adaptive Attention Temperature for Neural Machine Translation
- Code2Que: A Tool for Improving Question Titles from Mined Code Snippets in Stack Overflow
- Otem&Utem: Over- and Under-Translation Evaluation Metric for NMT
- Dual Past and Future for Neural Machine Translation