Training Tips for the Transformer Model
arXiv:1804.00247 · doi:10.2478/pralin-2018-0002
Abstract
This article describes our experiments in neural machine translation using the recent Tensor2Tensor framework and the Transformer sequence-to-sequence model (Vaswani et al., 2017). We examine some of the critical parameters that affect the final translation quality, memory usage, training stability and training time, concluding each experiment with a set of recommendations for fellow researchers. In addition to confirming the general mantra "more data and larger models", we address scaling to multiple GPUs and provide practical tips for improved training regarding batch size, learning rate, warmup steps, maximum sentence length and checkpoint averaging. We hope that our observations will allow others to get better results given their particular hardware and data constraints.
This is the version published in PBML (https://ufal.mff.cuni.cz/pbml/110/art-popel-bojar.pdf)
References in corpus (4)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- One weird trick for parallelizing convolutional neural networks
- On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima
Cited by in corpus (61)
- A Comparative Study on Transformer vs RNN in Speech Applications
- On the Variance of the Adaptive Learning Rate and Beyond
- Transformers in Time-series Analysis: A Tutorial
- Revisiting Few-sample BERT Fine-tuning
- Analyzing Multi-Head Self-Attention: Specialized Heads Do the Heavy Lifting, the Rest Can Be Pruned
- Data Diversification: A Simple Strategy For Neural Machine Translation
- Self-Paced Learning for Neural Machine Translation
- Prototypical Representation Learning for Relation Extraction
- BERT-of-Theseus: Compressing BERT by Progressive Module Replacing
- Constraint Translation Candidates: A Bridge between Neural Query Translation and Cross-lingual Information Retrieval
- Exploiting Neural Query Translation into Cross Lingual Information Retrieval
- A Neural Network Transformer Model for Composite Microstructure Homogenization
- Understanding the Difficulty of Training Transformers
- Passport: Improving Automated Formal Verification Using Identifiers
- A Comprehensive Survey of Grammar Error Correction
- SambaMixer: State of Health Prediction of Li-ion Batteries using Mamba State Space Models
- On Optimal Transformer Depth for Low-Resource Language Translation
- One Epoch Is All You Need
- Transformers: "The End of History" for NLP?
- DiscoFuse: A Large-Scale Dataset for Discourse-Based Sentence Fusion
- Announcing CzEng 2.0 Parallel Corpus with over 2 Gigawords
- Competence-based Curriculum Learning for Neural Machine Translation
- Exploring Benefits of Transfer Learning in Neural Machine Translation
- The Impact of LoRA Adapters on LLMs for Clinical Text Classification Under Computational and Data Constraints
- ESPnet-ST: All-in-One Speech Translation Toolkit
- Long-span language modeling for speech recognition
- CoMAE: A Multi-factor Hierarchical Framework for Empathetic Response Generation
- Disentangling Adaptive Gradient Methods from Learning Rates
- Predicting Actions to Help Predict Translations
- Better Sign Language Translation with STMC-Transformer
- Hotel2vec: Learning Attribute-Aware Hotel Embeddings with Self-Supervision
- Improving Transformer Performance for French Clinical Notes Classification Using Mixture of Experts on a Limited Dataset
- Evaluating Online Continual Learning with CALM
- Efficiency Metrics for Data-Driven Models: A Text Summarization Case Study
- Attend and Decode: 4D fMRI Task State Decoding Using Attention Models
- Retrosynthesis with Attention-Based NMT Model and Chemical Analysis of the "Wrong" Predictions
- Optimizing Deeper Transformers on Small Datasets
- Distributed Deep Learning in Open Collaborations
- GTNet:Guided Transformer Network for Detecting Human-Object Interactions
- Siamese Tracking with Lingual Object Constraints
- Learning Accurate Integer Transformer Machine-Translation Models
- A Novel Transformer-Based Self-Supervised Learning Method to Enhance Photoplethysmogram Signal Artifact Detection
- On the Distributional Properties of Adaptive Gradients
- Neural Machine Translation: A Review and Survey
- Learn Faster and Forget Slower via Fast and Stable Task Adaptation
- Making Asynchronous Stochastic Gradient Descent Work for Transformers
- On the Importance of Local Information in Transformer Based Models
- Improving Distinction between ASR Errors and Speech Disfluencies with Feature Space Interpolation
- ELITR Non-Native Speech Translation at IWSLT 2020
- In-training Matrix Factorization for Parameter-frugal Neural Machine Translation
- Context-Aware Monolingual Repair for Neural Machine Translation
- Machine Translation between Vietnamese and English: an Empirical Study
- CUNI System for the Building Educational Applications 2019 Shared Task: Grammatical Error Correction
- Learning Neural Models for Natural Language Processing in the Face of Distributional Shift
- Text Simplification for Comprehension-based Question-Answering
- Can the Transformer Be Used as a Drop-in Replacement for RNNs in Text-Generating GANs?
- Densifying Assumed-sparse Tensors: Improving Memory Efficiency and MPI Collective Performance during Tensor Accumulation for Parallelized Training of Neural Machine Translation Models
- Dynamically Adjusting Transformer Batch Size by Monitoring Gradient Direction Change
- Dynamic Curriculum Learning for Low-Resource Neural Machine Translation
- CVIT-MT Systems for WAT-2018
- Character-level Representations Improve DRS-based Semantic Parsing Even in the Age of BERT