Syntax-Directed Attention for Neural Machine Translation
arXiv:1711.04231
Abstract
Attention mechanism, including global attention and local attention, plays a key role in neural machine translation (NMT). Global attention attends to all source words for word prediction. In comparison, local attention selectively looks at fixed-window source words. However, alignment weights for the current target word often decrease to the left and right by linear distance centering on the aligned source position and neglect syntax-directed distance constraints. In this paper, we extend local attention with syntax-distance constraint, to focus on syntactically related source words with the predicted target word, thus learning a more effective context vector for word prediction. Moreover, we further propose a double context NMT architecture, which consists of a global context vector and a syntax-directed context vector over the global attention, to provide more translation performance for NMT from source representation. The experiments on the large-scale Chinese-to-English and English-to-Germen translation tasks show that the proposed approach achieves a substantial and significant improvement over the baseline system.
AAAI2018, revised version
Cited by in corpus (15)
- Attention in Natural Language Processing
- A Survey of Knowledge-Enhanced Text Generation
- SG-Net: Syntax-Guided Machine Reading Comprehension
- Head-Driven Phrase Structure Grammar Parsing on Penn Treebank
- Enhancing Machine Translation with Dependency-Aware Self-Attention
- Lattice-Based Transformer Encoder for Neural Machine Translation
- Discriminative Reasoning for Document-level Relation Extraction
- Tackling Graphical NLP problems with Graph Recurrent Networks
- Neural Machine Translation: A Review and Survey
- Open Vocabulary Learning for Neural Chinese Pinyin IME
- Effective Representation for Easy-First Dependency Parsing
- Regularized Context Gates on Transformer for Machine Translation
- Latent Part-of-Speech Sequences for Neural Machine Translation
- Structural Guidance for Transformer Language Models
- On the Relation between Syntactic Divergence and Zero-Shot Performance