Curriculum Learning and Minibatch Bucketing in Neural Machine Translation
arXiv:1707.09533 · doi:10.26615/978-954-452-049-6_050
Abstract
We examine the effects of particular orderings of sentence pairs on the on-line training of neural machine translation (NMT). We focus on two types of such orderings: (1) ensuring that each minibatch contains sentences similar in some aspect and (2) gradual inclusion of some sentence types as the training progresses (so called "curriculum learning"). In our English-to-Czech experiments, the internal homogeneity of minibatches has no effect on the training but some of our "curricula" achieve a small improvement over the baseline.
Accepted to RANLP 2017
References in corpus (6)
- Sequence to Sequence Learning with Neural Networks
- Natural Language Processing (almost) from Scratch
- Sequence Transduction with Recurrent Neural Networks
- Fast Domain Adaptation for Neural Machine Translation
- Automated Curriculum Learning for Neural Networks
- A comprehensive study of batch construction strategies for recurrent neural networks in MXNet
Cited by in corpus (35)
- Digital Twin-based Anomaly Detection with Curriculum Learning in Cyber-physical Systems
- Self-Paced Learning for Neural Machine Translation
- Recipes for Adapting Pre-trained Monolingual and Multilingual Models to Machine Translation
- MetaMT,a MetaLearning Method Leveraging Multiple Domain Data for Low Resource Machine Translation
- The Stability-Efficiency Dilemma: Investigating Sequence Length Warmup for Training GPT Models
- Competence-based Curriculum Learning for Neural Machine Translation
- Modeling Coherence for Discourse Neural Machine Translation
- A Survey on Curriculum Learning
- An Analytical Theory of Curriculum Learning in Teacher-Student Networks
- Norm-Based Curriculum Learning for Neural Machine Translation
- When Do Curricula Work?
- Data Rejuvenation: Exploiting Inactive Training Examples for Neural Machine Translation
- Curriculum Learning for Domain Adaptation in Neural Machine Translation
- Gradual Domain Adaptation in the Wild:When Intermediate Distributions are Absent
- Learning a Multi-Domain Curriculum for Neural Machine Translation
- Token-level Adaptive Training for Neural Machine Translation
- Dynamic Sentence Sampling for Efficient Training of Neural Machine Translation
- Assessing the Bilingual Knowledge Learned by Neural Machine Translation Models
- Curriculum Learning Strategies for IR: An Empirical Study on Conversation Response Ranking
- Reinforced Curriculum Learning on Pre-trained Neural Machine Translation Models
- Selective Knowledge Distillation for Neural Machine Translation
- Dynamically Composing Domain-Data Selection with Clean-Data Selection by "Co-Curricular Learning" for Neural Machine Translation
- Reinforcement Learning based Curriculum Optimization for Neural Machine Translation
- Curriculum Learning with Diversity for Supervised Computer Vision Tasks
- Statistical Measures For Defining Curriculum Scoring Function
- The LMU Munich System for the WMT 2020 Unsupervised Machine Translation Shared Task
- Self-Guided Curriculum Learning for Neural Machine Translation
- Competence-based Curriculum Learning for Multilingual Machine Translation
- Exploiting Curriculum Learning in Unsupervised Neural Machine Translation
- Bandits Don't Follow Rules: Balancing Multi-Facet Machine Translation with Multi-Armed Bandits
- Dynamic Curriculum Learning for Low-Resource Neural Machine Translation
- Data Ordering Patterns for Neural Machine Translation: An Empirical Study
- Does the Order of Training Samples Matter? Improving Neural Data-to-Text Generation with Curriculum Learning
- Token-wise Curriculum Learning for Neural Machine Translation
- Learning to Inpaint by Progressively Growing the Mask Regions