Rethinking Action Spaces for Reinforcement Learning in End-to-end Dialog Agents with Latent Variable Models
arXiv:1902.08858
Abstract
Defining action spaces for conversational agents and optimizing their decision-making process with reinforcement learning is an enduring challenge. Common practice has been to use handcrafted dialog acts, or the output vocabulary, e.g. in neural encoder decoders, as the action spaces. Both have their own limitations. This paper proposes a novel latent action framework that treats the action spaces of an end-to-end dialog agent as latent variables and develops unsupervised methods in order to induce its own action space from the data. Comprehensive experiments are conducted examining both continuous and discrete action types and two different optimization methods based on stochastic variational inference. Results show that the proposed latent actions achieve superior empirical performance improvement over previous word-level policy gradient methods on both DealOrNoDeal and MultiWoz dialogs. Our detailed analysis also provides insights about various latent variable approaches for policy learning and can serve as a foundation for developing better latent actions in future research.
Camera ready version for NAACL 2019 long paper
References in corpus (3)
Cited by in corpus (10)
- Task-Oriented Dialog Systems that Consider Multiple Appropriate Responses under the Same Context
- Robust Conversational AI with Grounded Text Generation
- Dialog without Dialog Data: Learning Visual Dialog Agents from VQA Data
- A Hybrid Task-Oriented Dialog System with Domain and Task Adaptive Pretraining
- OpenViDial 2.0: A Larger-Scale, Open-Domain Dialogue Generation Dataset with Visual Contexts
- Joint System-Wise Optimization for Pipeline Goal-Oriented Dialog System
- Resource Constrained Dialog Policy Learning via Differentiable Inductive Logic Programming
- MTSS: Learn from Multiple Domain Teachers and Become a Multi-domain Dialogue Expert
- Targeted Data Acquisition for Evolving Negotiation Agents
- When is it permissible for artificial intelligence to lie? A trust-based approach