Encoding Source Language with Convolutional Neural Network for Machine Translation
arXiv:1503.01838
Abstract
The recently proposed neural network joint model (NNJM) (Devlin et al., 2014) augments the n-gram target language model with a heuristically chosen source context window, achieving state-of-the-art performance in SMT. In this paper, we give a more systematic treatment by summarizing the relevant source information through a convolutional architecture guided by the target information. With different guiding signals during decoding, our specifically designed convolution+gating architectures can pinpoint the parts of a source sentence that are relevant to predicting a target word, and fuse them with the context of entire source sentence to form a unified representation. This representation, together with target language words, are fed to a deep neural network (DNN) to form a stronger NNJM. Experiments on two NIST Chinese-English translation tasks show that the proposed model can achieve significant improvements over the previous NNJM by up to +1.08 BLEU points on average
Accepted as a full paper at ACL 2015
References in corpus (3)
Cited by in corpus (12)
- One Model To Learn Them All
- Depthwise Separable Convolutions for Neural Machine Translation
- Question Answering and Question Generation as Dual Tasks
- A Survey of Deep Learning Techniques for Neural Machine Translation
- Deep Learning for Medical Image Segmentation
- Deconvolutional Paragraph Representation Learning
- Can Active Memory Replace Attention?
- Temporal Deformable Convolutional Encoder-Decoder Networks for Video Captioning
- Context-Dependent Translation Selection Using Convolutional Neural Network
- Exploring Different Dimensions of Attention for Uncertainty Detection
- A Sustainable Multi-modal Multi-layer Emotion-aware Service at the Edge
- Character-level Deep Conflation for Business Data Analytics