Non-Autoregressive Neural Machine Translation
arXiv:1711.02281
Abstract
Existing approaches to neural machine translation condition each output word on previously generated outputs. We introduce a model that avoids this autoregressive property and produces its outputs in parallel, allowing an order of magnitude lower latency during inference. Through knowledge distillation, the use of input token fertilities as a latent variable, and policy gradient fine-tuning, we achieve this at a cost of as little as 2.0 BLEU points relative to the autoregressive Transformer network used as a teacher. We demonstrate substantial cumulative improvements associated with each of the three aspects of our training strategy, and validate our approach on IWSLT 2016 English-German and two WMT language pairs. By sampling fertilities in parallel at inference time, our non-autoregressive model achieves near-state-of-the-art performance of 29.8 BLEU on WMT 2016 English-Romanian.
Accepted by ICLR 2018
References in corpus (3)
Cited by in corpus (108)
- FastSpeech: Fast, Robust and Controllable Text to Speech
- SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head Pruning
- Massively Multilingual Neural Machine Translation in the Wild: Findings and Challenges
- HAT: Hardware-Aware Transformers for Efficient Natural Language Processing
- Glow-TTS: A Generative Flow for Text-to-Speech via Monotonic Alignment Search
- Enable Deep Learning on Mobile Devices: Methods, Systems, and Applications
- Listen and Fill in the Missing Letters: Non-Autoregressive Transformer for Speech Recognition
- Paradigm Shift in Natural Language Processing
- Deep Learning Inference in Facebook Data Centers: Characterization, Performance Optimizations and Hardware Implications
- Non-Attentive Tacotron: Robust and Controllable Neural TTS Synthesis Including Unsupervised Duration Modeling
- Efficient Neural Audio Synthesis
- Non-Autoregressive Machine Translation with Disentangled Context Transformer
- ClariNet: Parallel Wave Generation in End-to-End Text-to-Speech
- Syntactically Supervised Transformers for Faster Neural Machine Translation
- Forecast Network-Wide Traffic States for Multiple Steps Ahead: A Deep Learning Approach Considering Dynamic Non-Local Spatial Correlation and Non-Stationary Temporal Dependency
- Text Compression-aided Transformer Encoding
- Challenges in Building Intelligent Open-domain Dialog Systems
- A Generalized Framework of Sequence Generation with Application to Undirected Sequence Models
- Incorporating BERT into Parallel Sequence Decoding with Adapters
- A CTC Alignment-based Non-autoregressive Transformer for End-to-end Automatic Speech Recognition
- Understanding Knowledge Distillation in Non-autoregressive Machine Translation
- Uncertainty Estimation in Autoregressive Structured Prediction
- Text-Conditioned Transformer for Automatic Pronunciation Error Detection
- Imitation Learning for Non-Autoregressive Neural Machine Translation
- Well Googled is Half Done: Multimodal Forecasting of New Fashion Product Sales with Image-based Google Trends
- Accelerating Transformer Inference for Translation via Parallel Decoding
- Non-autoregressive Transformer-based End-to-end ASR using BERT
- Insertion-based Decoding with automatically Inferred Generation Order
- Fast Image Caption Generation with Position Alignment
- Masked Non-Autoregressive Image Captioning
- BANG: Bridging Autoregressive and Non-autoregressive Generation with Large Scale Pretraining
- Accelerating Neural Transformer via an Average Attention Network
- Imputer: Sequence Modelling via Imputation and Dynamic Programming
- Neural Language Generation: Formulation, Methods, and Evaluation
- NAST: Non-Autoregressive Spatial-Temporal Transformer for Time Series Forecasting
- Parallel Tacotron: Non-Autoregressive and Controllable TTS
- Glancing Transformer for Non-Autoregressive Neural Machine Translation
- Contextualized Perturbation for Textual Adversarial Attack
- Non-Autoregressive Machine Translation with Auxiliary Regularization
- Probabilistic Forecasting with Temporal Convolutional Neural Network
- Strategies for Structuring Story Generation
- Non-Autoregressive Neural Machine Translation with Enhanced Decoder Input
- NAOMI: Non-Autoregressive Multiresolution Sequence Imputation
- Guiding Non-Autoregressive Neural Machine Translation Decoding with Reordering Information
- Latent-Variable Non-Autoregressive Neural Machine Translation with Deterministic Inference Using a Delta Posterior
- Cascaded Text Generation with Markov Transformers
- Hint-Based Training for Non-Autoregressive Machine Translation
- Non-Autoregressive Image Captioning with Counterfactuals-Critical Multi-Agent Learning
- Semi-Autoregressive Neural Machine Translation
- Deep Neural Machine Translation with Weakly-Recurrent Units
- Retrieving Sequential Information for Non-Autoregressive Neural Machine Translation
- An Efficient Transformer Decoder with Compressed Sub-layers
- Duplex Sequence-to-Sequence Learning for Reversible Machine Translation
- Spike-Triggered Non-Autoregressive Transformer for End-to-End Speech Recognition
- FastLTS: Non-Autoregressive End-to-End Unconstrained Lip-to-Speech Synthesis
- Neural Machine Translation: A Review of Methods, Resources, and Tools
- Fine-Tuning by Curriculum Learning for Non-Autoregressive Neural Machine Translation
- Improving Non-autoregressive Generation with Mixup Training
- Multi-Task Learning with Shared Encoder for Non-Autoregressive Machine Translation
- Hard but Robust, Easy but Sensitive: How Encoder and Decoder Perform in Neural Machine Translation
- XL-Editor: Post-editing Sentences with XLNet
- Exposing the Implicit Energy Networks behind Masked Language Models via Metropolis--Hastings
- Parallel Synthesis for Autoregressive Speech Generation
- Improved Mask-CTC for Non-Autoregressive End-to-End ASR
- Listen Attentively, and Spell Once: Whole Sentence Generation via a Non-Autoregressive Architecture for Low-Latency Speech Recognition
- End-to-End Human Object Interaction Detection with HOI Transformer
- Accelerating Transformer Decoding via a Hybrid of Self-attention and Recurrent Neural Network
- Neural Message Passing for Multi-Label Classification
- A Sequence-to-Set Network for Nested Named Entity Recognition
- Do sequence-to-sequence VAEs learn global features of sentences?
- Lifelong Language Knowledge Distillation
- Faster Re-translation Using Non-Autoregressive Model For Simultaneous Neural Machine Translation
- Improving Non-autoregressive Neural Machine Translation with Monolingual Data
- Multiresolution Transformer Networks: Recurrence is Not Essential for Modeling Hierarchical Structure
- Enriching Non-Autoregressive Transformer with Syntactic and SemanticStructures for Neural Machine Translation
- Computer Assisted Translation with Neural Quality Estimation and Automatic Post-Editing
- Orthros: Non-autoregressive End-to-end Speech Translation with Dual-decoder
- Rejuvenating Low-Frequency Words: Making the Most of Parallel Data in Non-Autoregressive Translation
- FastLR: Non-Autoregressive Lipreading Model with Integrate-and-Fire
- Neighbors Are Not Strangers: Improving Non-Autoregressive Translation under Low-Frequency Lexical Constraints
- Investigating the Reordering Capability in CTC-based Non-Autoregressive End-to-End Speech Translation
- Improved Variational Neural Machine Translation by Promoting Mutual Information
- Insertion-Based Modeling for End-to-End Automatic Speech Recognition
- Integrated Training for Sequence-to-Sequence Models Using Non-Autoregressive Transformer
- Learning to Recover from Multi-Modality Errors for Non-Autoregressive Neural Machine Translation
- Towards Variable-Length Textual Adversarial Attacks
- Straight to the Tree: Constituency Parsing with Neural Syntactic Distance
- Paraphrases as Foreign Languages in Multilingual Neural Machine Translation
- Semi-Autoregressive Transformer for Image Captioning
- Boundary and Context Aware Training for CIF-based Non-Autoregressive End-to-end ASR
- Neural Machine Translation: A Review and Survey
- Attending to Emotional Narratives
- Teaching a Massive Open Online Course on Natural Language Processing
- Streaming End-to-End ASR based on Blockwise Non-Autoregressive Models
- Towards Reinforcement Learning for Pivot-based Neural Machine Translation with Non-autoregressive Transformer
- Sequence Generation: From Both Sides to the Middle
- Non-autoregressive electron flow generation for reaction prediction
- Sentence-Permuted Paragraph Generation
- Autoregressive Knowledge Distillation through Imitation Learning
- Toward Streaming ASR with Non-Autoregressive Insertion-based Model
- Subword Language Model for Query Auto-Completion
- CASS-NAT: CTC Alignment-based Single Step Non-autoregressive Transformer for Speech Recognition
- Span Pointer Networks for Non-Autoregressive Task-Oriented Semantic Parsing
- GRET: Global Representation Enhanced Transformer
- Learning Energy-Based Approximate Inference Networks for Structured Applications in NLP
- Self-Guided Curriculum Learning for Neural Machine Translation
- How Does Distilled Data Complexity Impact the Quality and Confidence of Non-Autoregressive Machine Translation?
- Data-to-text Generation by Splicing Together Nearest Neighbors