Neural Sequence Model Training via -divergence Minimization
arXiv:1706.10031
Abstract
We propose a new neural sequence model training method in which the objective function is defined by -divergence. We demonstrate that the objective function generalizes the maximum-likelihood (ML)-based and reinforcement learning (RL)-based objective functions as special cases (i.e., ML corresponds to and RL to ). We also show that the gradient of the objective function can be considered a mixture of ML- and RL-based objective gradients. The experimental results of a machine translation task show that minimizing the objective function with outperforms , which corresponds to ML-based methods.
2017 ICML Workshop on Learning to Generate Natural Language (LGNL 2017)
References in corpus (5)
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation
- End-to-end Continuous Speech Recognition using Attention-based Recurrent NN: First Results
- Professor Forcing: A New Algorithm for Training Recurrent Networks
- An Actor-Critic Algorithm for Sequence Prediction