Learning Deep Transformer Models for Machine Translation
arXiv:1906.01787
Abstract
Transformer is the state-of-the-art model in recent machine translation evaluations. Two strands of research are promising to improve models of this kind: the first uses wide networks (a.k.a. Transformer-Big) and has been the de facto standard for the development of the Transformer system, and the other uses deeper language representation but faces the difficulty arising from learning deep networks. Here, we continue the line of research on the latter. We claim that a truly deep Transformer model can surpass the Transformer-Big counterpart by 1) proper use of layer normalization and 2) a novel way of passing the combination of previous layers to the next. On WMT'16 English- German, NIST OpenMT'12 Chinese-English and larger WMT'18 Chinese-English tasks, our deep system (30/25-layer encoder) outperforms the shallow Transformer-Big/Base baseline (6-layer encoder) by 0.4-2.4 BLEU points. As another bonus, the deep model is 1.6X smaller in size and 3X faster in training than Transformer-Big.
Accepted by ACL 2019
References in corpus (5)
- Sequence to Sequence Learning with Neural Networks
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- On the difficulty of training Recurrent Neural Networks
- Massive Exploration of Neural Machine Translation Architectures
- Multi-layer Representation Fusion for Neural Machine Translation
Cited by in corpus (14)
- A Survey of Deep Learning Techniques for Neural Machine Translation
- Accelerating Training of Transformer-Based Language Models with Progressive Layer Dropping
- Transformer Transducer: A Streamable Speech Recognition Model with Transformer Encoders and RNN-T Loss
- Centroid Transformers: Learning to Abstract with Attention
- Character-based NMT with Transformer
- The Volctrans GLAT System: Non-autoregressive Translation Meets WMT21
- Deep Ensembles on a Fixed Memory Budget: One Wide Network or Several Thinner Ones?
- Can Transformers Reason About Effects of Actions?
- Translating the Unseen? Yoruba-English MT in Low-Resource, Morphologically-Unmarked Settings
- Neural Machine Translation with Joint Representation
- Automated Query Reformulation for Efficient Search based on Query Logs From Stack Overflow
- Adversarial Joint Training with Self-Attention Mechanism for Robust End-to-End Speech Recognition
- Detecting and Understanding Generalization Barriers for Neural Machine Translation
- Translate Reverberated Speech to Anechoic Ones: Speech Dereverberation with BERT