Way Off-Policy Batch Deep Reinforcement Learning of Implicit Human Preferences in Dialog
arXiv:1907.00456
Abstract
Most deep reinforcement learning (RL) systems are not able to learn effectively from off-policy data, especially if they cannot explore online in the environment. These are critical shortcomings for applying RL to real-world problems where collecting data is expensive, and models must be tested offline before being deployed to interact with the environment -- e.g. systems that learn from human interaction. Thus, we develop a novel class of off-policy batch RL algorithms, which are able to effectively learn offline, without exploring, from a fixed batch of human interaction data. We leverage models pre-trained on data as a strong prior, and use KL-control to penalize divergence from this prior during RL training. We also use dropout-based uncertainty estimates to lower bound the target Q-values as a more efficient alternative to Double Q-Learning. The algorithms are tested on the problem of open-domain dialog generation -- a challenging reinforcement learning problem with a 20,000-dimensional action space. Using our Way Off-Policy algorithm, we can extract multiple different reward functions post-hoc from collected human interaction data, and learn effectively from all of these. We test the real-world generalization of these systems by deploying them live to converse with humans in an open-domain setting, and demonstrate that our algorithm achieves significant improvements over prior methods in off-policy batch RL.
References in corpus (6)
- Distilling the Knowledge in a Neural Network
- Chameleons in imagined conversations: A new approach to understanding coordination of linguistic style in dialogs
- Uncertainty-Aware Reinforcement Learning for Collision Avoidance
- A Deep Reinforcement Learning Chatbot
- Dialogue Learning With Human-In-The-Loop
- Off-Policy Policy Gradient with State Distribution Correction
Cited by in corpus (26)
- Conservative Q-Learning for Offline Reinforcement Learning
- Behavior Regularized Offline Reinforcement Learning
- Benchmarking Batch Deep Reinforcement Learning Algorithms
- Recursively Summarizing Books with Human Feedback
- Deployment-Efficient Reinforcement Learning via Model-Based Offline Optimization
- COG: Connecting New Skills to Past Experience with Offline Reinforcement Learning
- Offline Reinforcement Learning with Fisher Divergence Critic Regularization
- Neural Language Generation: Formulation, Methods, and Evaluation
- Representation Matters: Offline Pretraining for Sequential Decision Making
- Uncertainty Weighted Actor-Critic for Offline Reinforcement Learning
- Believe What You See: Implicit Constraint Approach for Offline Multi-Agent Reinforcement Learning
- Risk-Averse Offline Reinforcement Learning
- Provable Benefits of Actor-Critic Methods for Offline Reinforcement Learning
- Behavior Priors for Efficient Reinforcement Learning
- PLAS: Latent Action Space for Offline Reinforcement Learning
- Behavioral Priors and Dynamics Models: Improving Performance and Domain Transfer in Offline RL
- Batch-Constrained Distributional Reinforcement Learning for Session-based Recommendation
- A Workflow for Offline Model-Free Robotic Reinforcement Learning
- Document-editing Assistants and Model-based Reinforcement Learning as a Path to Conversational AI
- Regularized Behavior Value Estimation
- Conservative Data Sharing for Multi-Task Offline Reinforcement Learning
- On the Optimality of Batch Policy Optimization Algorithms
- Reducing Conservativeness Oriented Offline Reinforcement Learning
- Learning Natural Language Generation from Scratch
- Off-Policy Self-Critical Training for Transformer in Visual Paragraph Generation
- Offline Inverse Reinforcement Learning