Personalizing a Dialogue System with Transfer Reinforcement Learning
arXiv:1610.02891
Abstract
It is difficult to train a personalized task-oriented dialogue system because the data collected from each individual is often insufficient. Personalized dialogue systems trained on a small dataset can overfit and make it difficult to adapt to different user needs. One way to solve this problem is to consider a collection of multiple users' data as a source domain and an individual user's data as a target domain, and to perform a transfer learning from the source to the target domain. By following this idea, we propose "PETAL"(PErsonalized Task-oriented diALogue), a transfer-learning framework based on POMDP to learn a personalized dialogue system. The system first learns common dialogue knowledge from the source domain and then adapts this knowledge to the target user. This framework can avoid the negative transfer problem by considering differences between source and target users. The policy in the personalized POMDP can learn to choose different actions appropriately for different users. Experimental results on a real-world coffee-shopping data and simulation data show that our personalized dialogue system can choose different optimal actions for different users, and thus effectively improve the dialogue quality under the personalized setting.
References in corpus (5)
- Deep Reinforcement Learning for Dialogue Generation
- A Personalized System for Conversational Recommendations
- A Network-based End-to-End Trainable Task-oriented Dialogue System
- End-to-end LSTM-based dialog control optimized with supervised and reinforcement learning
- deltaBLEU: A Discriminative Metric for Generation Tasks with Intrinsically Diverse Targets
Cited by in corpus (6)
- A Survey on Dialogue Systems: Recent Advances and New Frontiers
- MALA: Cross-Domain Dialogue Generation with Action Learning
- Learning to Transfer
- All by Myself: Learning Individualized Competitive Behaviour with a Contrastive Reinforcement Learning optimization
- Generating Multiple Diverse Responses for Short-Text Conversation
- Cross-domain Dialogue Policy Transfer via Simultaneous Speech-act and Slot Alignment