The Evolved Transformer
arXiv:1901.11117
Abstract
Recent works have highlighted the strength of the Transformer architecture on sequence tasks while, at the same time, neural architecture search (NAS) has begun to outperform human-designed models. Our goal is to apply NAS to search for a better alternative to the Transformer. We first construct a large search space inspired by the recent advances in feed-forward sequence models and then run evolutionary architecture search with warm starting by seeding our initial population with the Transformer. To directly search on the computationally expensive WMT 2014 English-German translation task, we develop the Progressive Dynamic Hurdles method, which allows us to dynamically allocate more resources to more promising candidate models. The architecture found in our experiments -- the Evolved Transformer -- demonstrates consistent improvement over the Transformer on four well-established language tasks: WMT 2014 English-German, WMT 2014 English-French, WMT 2014 English-Czech and LM1B. At a big model size, the Evolved Transformer establishes a new state-of-the-art BLEU score of 29.8 on WMT'14 English-German; at smaller sizes, it achieves the same quality as the original "big" Transformer with 37.6% less parameters and outperforms the Transformer by 0.7 BLEU at a mobile-friendly model size of 7M parameters.
ICML version with SOTA results
Cited by in corpus (19)
- Towards a Human-like Open-Domain Chatbot
- A Survey on Neural Architecture Search
- Incorporating BERT into Neural Machine Translation
- Funnel-Transformer: Filtering out Sequential Redundancy for Efficient Language Processing
- A Survey of Deep Learning Techniques for Neural Machine Translation
- MUSE: Parallel Multi-Scale Attention for Sequence to Sequence Learning
- Weight-Sharing Neural Architecture Search: A Battle to Shrink the Optimization Gap
- Neural Language Generation: Formulation, Methods, and Evaluation
- Multi-branch Attentive Transformer
- Finding Fast Transformers: One-Shot Neural Architecture Search by Component Composition
- Proof of Learning (PoLe): Empowering Machine Learning with Consensus Building on Blockchains
- Self-Segregating and Coordinated-Segregating Transformer for Focused Deep Multi-Modular Network for Visual Question Answering
- AutoRC: Improving BERT Based Relation Classification Models via Architecture Search
- SACT: Self-Aware Multi-Space Feature Composition Transformer for Multinomial Attention for Video Captioning
- Multi-Pass Transformer for Machine Translation
- You May Not Need Order in Time Series Forecasting
- Evolving Multi-Resolution Pooling CNN for Monaural Singing Voice Separation
- Learning Architectures from an Extended Search Space for Language Modeling
- Two-Headed Monster And Crossed Co-Attention Networks