A Convolutional Encoder Model for Neural Machine Translation
arXiv:1611.02344
Abstract
The prevalent approach to neural machine translation relies on bi-directional LSTMs to encode the source sentence. In this paper we present a faster and simpler architecture based on a succession of convolutional layers. This allows to encode the entire source sentence simultaneously compared to recurrent networks for which computation is constrained by temporal dependencies. On WMT'16 English-Romanian translation we achieve competitive accuracy to the state-of-the-art and we outperform several recently published results on the WMT'15 English-German task. Our models obtain almost the same accuracy as a very deep LSTM setup on WMT'14 English-French translation. Our convolutional encoder speeds up CPU decoding by more than two times at the same or higher accuracy as a strong bi-directional LSTM baseline.
13 pages
References in corpus (13)
- Adam: A Method for Stochastic Optimization
- Sequence to Sequence Learning with Neural Networks
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- Natural Language Processing (almost) from Scratch
- Sequence Level Training with Recurrent Neural Networks
- Is Neural Machine Translation Ready for Deployment? A Case Study on 30 Translation Directions
- On Using Very Large Target Vocabulary for Neural Machine Translation
- Encoding Source Language with Convolutional Neural Network for Machine Translation
- Vocabulary Selection Strategies for Neural Machine Translation
- Deep Recurrent Models with Fast-Forward Connections for Neural Machine Translation
- Neural Machine Translation with Recurrent Attention Modeling
- Context-Dependent Translation Selection Using Convolutional Neural Network
- Vocabulary Manipulation for Neural Machine Translation
Cited by in corpus (36)
- An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling
- Convolutional Sequence to Sequence Learning
- GCNet: Non-local Networks Meet Squeeze-Excitation Networks and Beyond
- Deformable ConvNets v2: More Deformable, Better Results
- Multi-level Convolutional Autoencoder Networks for Parametric Prediction of Spatio-temporal Dynamics
- Weighted Transformer Network for Machine Translation
- Sharp Minima Can Generalize For Deep Nets
- A Survey of Deep Learning Techniques for Neural Machine Translation
- R-Transformer: Recurrent Neural Network Enhanced Transformer
- Massive Exploration of Neural Machine Translation Architectures
- Relation Networks for Object Detection
- GeniePath: Graph Neural Networks with Adaptive Receptive Paths
- Dual-Primal Graph Convolutional Networks
- Attacking Visual Language Grounding with Adversarial Examples: A Case Study on Neural Image Captioning
- Link Prediction via Graph Attention Network
- Learning Generic Sentence Representations Using Convolutional Neural Networks
- Trainable Greedy Decoding for Neural Machine Translation
- A Comprehensive Survey of Deep Learning for Image Captioning
- Acquiring Knowledge from Pre-trained Model to Neural Machine Translation
- BLK-REW: A Unified Block-based DNN Pruning Framework using Reweighted Regularization Method
- Hint-Based Training for Non-Autoregressive Machine Translation
- Neural Classification of Malicious Scripts: A study with JavaScript and VBScript
- Data Augmentation for Skin Lesion using Self-Attention based Progressive Generative Adversarial Network
- Improving Neural Machine Translation with Pre-trained Representation
- Learning Efficient Lexically-Constrained Neural Machine Translation with External Memory
- Refining Source Representations with Relation Networks for Neural Machine Translation
- PAM:Point-wise Attention Module for 6D Object Pose Estimation
- Analyzing and Interpreting Convolutional Neural Networks in NLP
- Fine Grained Human Evaluation for English-to-Chinese Machine Translation: A Case Study on Scientific Text
- A Hierarchical Neural Network for Sequence-to-Sequences Learning
- Global Context Networks
- SocialGrid: A TCN-enhanced Method for Online Discussion Forecasting
- GRET: Global Representation Enhanced Transformer
- Logographic Subword Model for Neural Machine Translation
- Sentence-wise Smooth Regularization for Sequence to Sequence Learning
- Lookup subnet based Spatial Graph Convolutional neural Network