Neural Machine Translation in Linear Time
arXiv:1610.10099
Abstract
We present a novel neural network for processing sequences. The ByteNet is a one-dimensional convolutional neural network that is composed of two parts, one to encode the source sequence and the other to decode the target sequence. The two network parts are connected by stacking the decoder on top of the encoder and preserving the temporal resolution of the sequences. To address the differing lengths of the source and the target, we introduce an efficient mechanism by which the decoder is dynamically unfolded over the representation of the encoder. The ByteNet uses dilation in the convolutional layers to increase its receptive field. The resulting network has two core properties: it runs in time that is linear in the length of the sequences and it sidesteps the need for excessive memorization. The ByteNet decoder attains state-of-the-art performance on character-level language modelling and outperforms the previous best results obtained with recurrent networks. The ByteNet also achieves state-of-the-art performance on character-to-character machine translation on the English-to-German WMT translation task, surpassing comparable neural translation models that are based on recurrent networks with attentional pooling and run in quadratic time. We find that the latent alignment structure contained in the representations reflects the expected alignment between the tokens.
9 pages
References in corpus (1)
Cited by in corpus (57)
- Quasi-Recurrent Neural Networks
- One Model To Learn Them All
- Depthwise Separable Convolutions for Neural Machine Translation
- The Reversible Residual Network: Backpropagation Without Storing Activations
- Weighted Transformer Network for Machine Translation
- Neural Machine Translation and Sequence-to-sequence Models: A Tutorial
- Generating and designing DNA with deep generative models
- eXpose: A Character-Level Convolutional Neural Network with Embeddings For Detecting Malicious URLs, File Paths and Registry Keys
- A Survey of Deep Learning Techniques for Neural Machine Translation
- PixelCNN++: Improving the PixelCNN with Discretized Logistic Mixture Likelihood and Other Modifications
- Graph WaveNet for Deep Spatial-Temporal Graph Modeling
- Dynamic Evaluation of Neural Sequence Models
- BP-Transformer: Modelling Long-Range Context via Binary Partitioning
- Compressive Transformers for Long-Range Sequence Modelling
- Parallel Multiscale Autoregressive Density Estimation
- A Hybrid Convolutional Variational Autoencoder for Text Generation
- Neural Text Generation: A Practical Guide
- Unsupervised Cipher Cracking Using Discrete GANs
- Explainable artificial intelligence model to predict acute critical illness from electronic health records
- FastWave: Accelerating Autoregressive Convolutional Neural Networks on FPGA
- Neural Language Generation: Formulation, Methods, and Evaluation
- Extreme Multi-Label Legal Text Classification: A case study in EU Legislation
- Neural message passing for joint paratope-epitope prediction
- Non-Autoregressive Neural Machine Translation with Enhanced Decoder Input
- Trainable Greedy Decoding for Neural Machine Translation
- Temporal Deformable Convolutional Encoder-Decoder Networks for Video Captioning
- Hard-Coded Gaussian Attention for Neural Machine Translation
- How to Hallucinate Functional Proteins
- Sequential Deep Learning for Credit Risk Monitoring with Tabular Financial Data
- Finite Volume Neural Network: Modeling Subsurface Contaminant Transport
- LAVA NAT: A Non-Autoregressive Translation Model with Look-Around Decoding and Vocabulary Attention
- 1D Convolutional Neural Network Models for Sleep Arousal Detection
- Joint Echo Cancellation and Noise Suppression based on Cascaded Magnitude and Complex Mask Estimation
- Neural Machine Translation: A Review of Methods, Resources, and Tools
- CNN Is All You Need
- Anytime Sampling for Autoregressive Models via Ordered Autoencoding
- Spatial PixelCNN: Generating Images from Patches
- Hierarchical Sequence to Sequence Voice Conversion with Limited Data
- Discovery of Natural Language Concepts in Individual Units of CNNs
- User-specific Adaptive Fine-tuning for Cross-domain Recommendations
- Infusing Sequential Information into Conditional Masked Translation Model with Self-Review Mechanism
- A Feasible Framework for Arbitrary-Shaped Scene Text Recognition
- Enhanced 3D Human Pose Estimation from Videos by using Attention-Based Neural Network with Dilated Convolutions
- Recurrent Point Review Models
- Recurrent Neural Network from Adder's Perspective: Carry-lookahead RNN
- Cut-Based Graph Learning Networks to Discover Compositional Structure of Sequential Video Data
- Neural Contract Element Extraction Revisited: Letters from Sesame Street
- Upgrading the Newsroom: An Automated Image Selection System for News Articles
- Formant Tracking Using Dilated Convolutional Networks Through Dense Connection with Gating Mechanism
- 3D human pose estimation with adaptive receptive fields and dilated temporal convolutions
- Sentence-wise Smooth Regularization for Sequence to Sequence Learning
- Back to the Future: Joint Aware Temporal Deep Learning 3D Human Pose Estimation
- Temporally Folded Convolutional Neural Networks for Sequence Forecasting
- Water Supply Prediction Based on Initialized Attention Residual Network
- Understanding Feature Selection and Feature Memorization in Recurrent Neural Networks
- AI-lead Court Debate Case Investigation
- Developing neural machine translation models for Hungarian-English