paper

Neural Sequence Model Training via -divergence Minimization

arXiv:1706.10031

Abstract

We propose a new neural sequence model training method in which the objective function is defined by -divergence. We demonstrate that the objective function generalizes the maximum-likelihood (ML)-based and reinforcement learning (RL)-based objective functions as special cases (i.e., ML corresponds to and RL to ). We also show that the gradient of the objective function can be considered a mixture of ML- and RL-based objective gradients. The experimental results of a machine translation task show that minimizing the objective function with outperforms , which corresponds to ML-based methods.

2017 ICML Workshop on Learning to Generate Natural Language (LGNL 2017)

References in corpus (5)

Neural Sequence Model Training via $α$-divergence Minimization · wovepaper