Discourse-Aware Neural Rewards for Coherent Text Generation
arXiv:1805.03766
Abstract
In this paper, we investigate the use of discourse-aware rewards with reinforcement learning to guide a model to generate long, coherent text. In particular, we propose to learn neural rewards to model cross-sentence ordering as a means to approximate desired discourse structure. Empirical results demonstrate that a generator trained with the learned reward produces more coherent and less repetitive text than models trained with cross-entropy or with reinforcement learning with commonly used scores as rewards.
NAACL 2018
References in corpus (7)
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- A Deep Reinforced Model for Abstractive Summarization
- Search-based Structured Prediction
- Deep Reinforcement Learning-based Image Captioning with Embedding Reward
- Challenges in Data-to-Document Generation
- Simulating Action Dynamics with Neural Process Networks
- Paraphrase Generation with Deep Reinforcement Learning