Towards End-to-End Learning for Dialog State Tracking and Management using Deep Reinforcement Learning
arXiv:1606.02560
Abstract
This paper presents an end-to-end framework for task-oriented dialog systems using a variant of Deep Recurrent Q-Networks (DRQN). The model is able to interface with a relational database and jointly learn policies for both language understanding and dialog strategy. Moreover, we propose a hybrid algorithm that combines the strength of reinforcement learning and supervised learning to achieve faster learning speed. We evaluated the proposed model on a 20 Question Game conversational game simulator. Results show that the proposed method outperforms the modular-based baseline and learns a distributed representation of the latent dialog state.
In proceeding of SIGDIAL 2016. Added changes based-on peer review, including: 1. Added references, 2. fixed typos in text and figures, 3. added minor change to introduction
References in corpus (3)
Cited by in corpus (49)
- Deep Reinforcement Learning: An Overview
- A Deep Reinforcement Learning Chatbot
- Learning Discourse-level Diversity for Neural Dialog Models using Conditional Variational Autoencoders
- A User Simulator for Task-Completion Dialogues
- Learn What Not to Learn: Action Elimination with Deep Reinforcement Learning
- SOLOIST: Building Task Bots at Scale with Transfer Learning and Machine Teaching
- An End-to-End Trainable Neural Network Model with Belief Tracking for Task-Oriented Dialog
- Towards End-to-End Reinforcement Learning of Dialogue Agents for Information Access
- Composite Task-Completion Dialogue Policy Learning via Hierarchical Deep Reinforcement Learning
- A Benchmarking Environment for Reinforcement Learning Based Task Oriented Dialogue Management
- Investigation of Language Understanding Impact for Reinforcement Learning Based Dialogue Systems
- Flexible and Scalable State Tracking Framework for Goal-Oriented Dialogue Systems
- Dialogue Learning with Human Teaching and Feedback in End-to-End Trainable Task-Oriented Dialogue Systems
- An End-to-end Approach for Handling Unknown Slot Values in Dialogue State Tracking
- Learning Robust Dialog Policies in Noisy Environments
- Zero-Shot Dialog Generation with Cross-Domain Latent Actions
- Learning End-to-End Goal-Oriented Dialog with Maximal User Task Success and Minimal Human Agent Use
- Iterative Policy Learning in End-to-End Trainable Task-Oriented Neural Dialog Models
- User Modeling for Task Oriented Dialogues
- Multimodal Hierarchical Reinforcement Learning Policy for Task-Oriented Visual Dialog
- Generative Encoder-Decoder Models for Task-Oriented Spoken Dialog Systems with Chatting Capability
- Investigation of Error Simulation Techniques for Learning Dialog Policies for Conversational Error Recovery
- Discriminative Deep Dyna-Q: Robust Planning for Dialogue Policy Learning
- Sentiment Adaptive End-to-End Dialog Systems
- Deep Active Learning for Dialogue Generation
- The Technological Gap Between Virtual Assistants and Recommendation Systems
- A Deep Reinforcement Learning Approach for Traffic Signal Control Optimization
- Prototypical Q Networks for Automatic Conversational Diagnosis and Few-Shot New Disease Adaption
- Report from the NSF Future Directions Workshop, Toward User-Oriented Agents: Research Directions and Challenges
- Towards Learning Transferable Conversational Skills using Multi-dimensional Dialogue Modelling
- Beyond task success: A closer look at jointly learning to see, ask, and GuessWhat
- Multi-Task Learning for Situated Multi-Domain End-to-End Dialogue Systems
- Joint System-Wise Optimization for Pipeline Goal-Oriented Dialog System
- Towards Personalized Dialog Policies for Conversational Skill Discovery
- Interactive Teaching for Conversational AI
- DRL: Deep Reinforcement Learning for Intelligent Robot Control -- Concept, Literature, and Future
- Adversarial Advantage Actor-Critic Model for Task-Completion Dialogue Policy Learning
- Dialogue Act Classification in Group Chats with DAG-LSTMs
- Adversarial Learning of Task-Oriented Neural Dialog Models
- Domain Transfer in Dialogue Systems without Turn-Level Supervision
- Learning Goal-oriented Dialogue Policy with Opposite Agent Awareness
- Building Advanced Dialogue Managers for Goal-Oriented Dialogue Systems
- Every time I fire a conversational designer, the performance of the dialog system goes down
- Learning to Ask Medical Questions using Reinforcement Learning
- Variational Reward Estimator Bottleneck: Learning Robust Reward Estimator for Multi-Domain Task-Oriented Dialog
- Improving Search through A3C Reinforcement Learning based Conversational Agent
- Hybrid Supervised Reinforced Model for Dialogue Systems
- Incremental Learning from Scratch for Task-Oriented Dialogue Systems
- Context-Aware Dialog Re-Ranking for Task-Oriented Dialog Systems