Boosting Neural Machine Translation
arXiv:1612.06138
Abstract
Training efficiency is one of the main problems for Neural Machine Translation (NMT). Deep networks need for very large data as well as many training iterations to achieve state-of-the-art performance. This results in very high computation cost, slowing down research and industrialisation. In this paper, we propose to alleviate this problem with several training methods based on data boosting and bootstrap with no modifications to the neural network. It imitates the learning process of humans, which typically spend more time when learning "difficult" concepts than easier ones. We experiment on an English-French translation task showing accuracy improvements of up to 1.63 BLEU while saving 20% of training time.
published in IJCNLP 2017
References in corpus (7)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Sequence to Sequence Learning with Neural Networks
- Convolutional Sequence to Sequence Learning
- Quantized Neural Networks: Training Neural Networks with Low Precision Weights and Activations
- Neural Machine Translation in Linear Time
- Factorization tricks for LSTM networks
- Deep Recurrent Models with Fast-Forward Connections for Neural Machine Translation
Cited by in corpus (8)
- An Empirical Exploration of Curriculum Learning for Neural Machine Translation
- Norm-Based Curriculum Learning for Neural Machine Translation
- Dynamic Sentence Sampling for Efficient Training of Neural Machine Translation
- Neural Machine Translation: A Review and Survey
- Curriculum Learning Strategies for IR: An Empirical Study on Conversation Response Ranking
- Data Ordering Patterns for Neural Machine Translation: An Empirical Study
- LSTMs Compose (and Learn) Bottom-Up
- Exploiting Curriculum Learning in Unsupervised Neural Machine Translation