BBQ-Networks: Efficient Exploration in Deep Reinforcement Learning for Task-Oriented Dialogue Systems
arXiv:1608.05081
Abstract
We present a new algorithm that significantly improves the efficiency of exploration for deep Q-learning agents in dialogue systems. Our agents explore via Thompson sampling, drawing Monte Carlo samples from a Bayes-by-Backprop neural network. Our algorithm learns much faster than common exploration strategies such as -greedy, Boltzmann, bootstrapping, and intrinsic-reward-based ones. Additionally, we show that spiking the replay buffer with experiences from just a few successful episodes can make Q-learning feasible when it might otherwise fail.
13 pages, 9 figures
Cited by in corpus (35)
- Modeling Multi-turn Conversation with Deep Utterance Aggregation
- Estimating Risk and Uncertainty in Deep Reinforcement Learning
- Challenges in Building Intelligent Open-domain Dialog Systems
- Neural Thompson Sampling
- Agnostic Q-learning with Function Approximation in Deterministic Systems: Tight Bounds on Approximation Error and Sample Complexity
- Uncertainty Weighted Actor-Critic for Offline Reinforcement Learning
- Robust Conversational AI with Grounded Text Generation
- Variational inference for the multi-armed contextual bandit
- Guided Dialog Policy Learning without Adversarial Learning in the Loop
- Randomized Exploration in Generalized Linear Bandits
- Show Us the Way: Learning to Manage Dialog from Demonstrations
- Meta Dialogue Policy Learning
- Bayesian Curiosity for Efficient Exploration in Reinforcement Learning
- Emotional Neural Language Generation Grounded in Situational Contexts
- Empirical Bayes Regret Minimization
- End-to-End Knowledge-Routed Relational Dialogue System for Automatic Diagnosis
- Multi-facet Contextual Bandits: A Neural Network Perspective
- Document-editing Assistants and Model-based Reinforcement Learning as a Path to Conversational AI
- What Should I Ask? Using Conversationally Informative Rewards for Goal-Oriented Visual Dialog
- Rethinking Supervised Learning and Reinforcement Learning in Task-Oriented Dialogue Systems
- An Overview of Natural Language State Representation for Reinforcement Learning
- Predict-then-Decide: A Predictive Approach for Wait or Answer Task in Dialogue Systems
- Provably Efficient Exploration for Reinforcement Learning Using Unsupervised Learning
- DialogAct2Vec: Towards End-to-End Dialogue Agent by Multi-Task Representation Learning
- Towards End-to-End Learning for Efficient Dialogue Agent by Modeling Looking-ahead Ability
- Sample-Efficient Model-based Actor-Critic for an Interactive Dialogue Task
- High-Quality Diversification for Task-Oriented Dialogue Systems
- Dynamic Knowledge Routing Network For Target-Guided Open-Domain Conversation
- "Wait, I'm Still Talking!" Predicting the Dialogue Interaction Behavior Using Imagine-Then-Arbitrate Model
- Learning Goal-oriented Dialogue Policy with Opposite Agent Awareness
- Where Do Human Heuristics Come From?
- Generative Dialog Policy for Task-oriented Dialog Systems
- Hybrid Supervised Reinforced Model for Dialogue Systems
- Amortized Variational Deep Q Network
- Incremental Learning from Scratch for Task-Oriented Dialogue Systems