Learning a Simple and Effective Model for Multi-turn Response Generation with Auxiliary Tasks
arXiv:2004.01972
Abstract
We study multi-turn response generation for open-domain dialogues. The existing state-of-the-art addresses the problem with deep neural architectures. While these models improved response quality, their complexity also hinders the application of the models in real systems. In this work, we pursue a model that has a simple structure yet can effectively leverage conversation contexts for response generation. To this end, we propose four auxiliary tasks including word order recovery, utterance order recovery, masked word recovery, and masked utterance recovery, and optimize the objectives of these tasks together with maximizing the likelihood of generation. By this means, the auxiliary tasks that relate to context understanding can guide the learning of the generation model to achieve a better local optimum. Empirical studies with three benchmarks indicate that our model can significantly outperform state-of-the-art generation models in terms of response quality on both automatic evaluation and human judgment, and at the same time enjoys a much faster decoding process.
References in corpus (8)
- Sequence to Sequence Learning with Neural Networks
- A Neural Conversational Model
- Unified Language Model Pre-training for Natural Language Understanding and Generation
- Wizard of Wikipedia: Knowledge-Powered Conversational agents
- Topic Aware Neural Response Generation
- Order Matters: Sequence to sequence for sets
- Hierarchical Recurrent Attention Network for Response Generation
- Improving Multi-turn Dialogue Modelling with Utterance ReWriter
Cited by in corpus (5)
- Learning an Effective Context-Response Matching Model with Self-Supervised Tasks for Retrieval-based Dialogues
- Towards Standard Criteria for human evaluation of Chatbots: A Survey
- ProphetNet-X: Large-Scale Pre-training Models for English, Chinese, Multi-lingual, Dialog, and Code Generation
- A Speaker-aware Parallel Hierarchical Attentive Encoder-Decoder Model for Multi-turn Dialogue Generation
- Advances in Multi-turn Dialogue Comprehension: A Survey