Deep Encoder, Shallow Decoder: Reevaluating Non-autoregressive Machine Translation
arXiv:2006.10369
Abstract
Much recent effort has been invested in non-autoregressive neural machine translation, which appears to be an efficient alternative to state-of-the-art autoregressive machine translation on modern GPUs. In contrast to the latter, where generation is sequential, the former allows generation to be parallelized across target token positions. Some of the latest non-autoregressive models have achieved impressive translation quality-speed tradeoffs compared to autoregressive baselines. In this work, we reexamine this tradeoff and argue that autoregressive baselines can be substantially sped up without loss in accuracy. Specifically, we study autoregressive models with encoders and decoders of varied depths. Our extensive experiments show that given a sufficiently deep encoder, a single-layer autoregressive decoder can substantially outperform strong non-autoregressive models with comparable inference speed. We show that the speed disadvantage for autoregressive baselines compared to non-autoregressive methods has been overestimated in three aspects: suboptimal layer allocation, insufficient speed measurement, and lack of knowledge distillation. Our results establish a new protocol for future research toward fast, accurate machine translation. Our code is available at https://github.com/jungokasai/deep-shallow.
ICLR 2021 Final Version
References in corpus (15)
- Distilling the Knowledge in a Neural Network
- Sequence to Sequence Learning with Neural Networks
- Multilingual Denoising Pre-training for Neural Machine Translation
- Reformer: The Efficient Transformer
- Pay Less Attention with Lightweight and Dynamic Convolutions
- Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention
- Tying Word Vectors and Word Classifiers: A Loss Framework for Language Modeling
- Insertion Transformer: Flexible Sequence Generation via Insertion Operations
- Aligned Cross Entropy for Non-Autoregressive Machine Translation
- Non-Autoregressive Machine Translation with Disentangled Context Transformer
- KERMIT: Generative Insertion-Based Modeling for Sequences
- A Generalized Framework of Sequence Generation with Application to Undirected Sequence Models
- Minimizing the Bag-of-Ngrams Difference for Non-Autoregressive Neural Machine Translation
- Faster Transformer Decoding: N-gram Masked Self-Attention
- LAVA NAT: A Non-Autoregressive Translation Model with Look-Around Decoding and Vocabulary Attention
Cited by in corpus (10)
- CPT: A Pre-Trained Unbalanced Transformer for Both Chinese Language Understanding and Generation
- Charformer: Fast Character Transformers via Gradient-based Subword Tokenization
- Scalable and Efficient MoE Training for Multitask Multilingual Models
- Scaling Laws for Neural Machine Translation
- Non-autoregressive End-to-end Speech Translation with Parallel Autoregressive Rescoring
- Probing Word Translations in the Transformer and Trading Decoder for Encoder Layers
- NVIDIA NeMo Neural Machine Translation Systems for English-German and English-Russian News and Biomedical Tasks at WMT21
- Diversifying Dialog Generation via Adaptive Label Smoothing
- Sentence Bottleneck Autoencoders from Transformer Language Models
- AligNART: Non-autoregressive Neural Machine Translation by Jointly Learning to Estimate Alignment and Translate