Recurrent Reinforcement Learning: A Hybrid Approach
arXiv:1509.03044
Abstract
Successful applications of reinforcement learning in real-world problems often require dealing with partially observable states. It is in general very challenging to construct and infer hidden states as they often depend on the agent's entire interaction history and may require substantial domain knowledge. In this work, we investigate a deep-learning approach to learning the representation of states in partially observable tasks, with minimal prior knowledge of the domain. In particular, we propose a new family of hybrid models that combines the strength of both supervised learning (SL) and reinforcement learning (RL), trained in a joint fashion: The SL component can be a recurrent neural networks (RNN) or its long short-term memory (LSTM) version, which is equipped with the desired property of being able to capture long-term dependency on history, thus providing an effective way of learning the representation of hidden states. The RL component is a deep Q-network (DQN) that learns to optimize the control for maximizing long-term rewards. Extensive experiments in a direct mailing campaign problem demonstrate the effectiveness and advantages of the proposed approach, which performs the best among a set of previous state-of-the-art methods.
11 pages, 6 figures
References in corpus (2)
Cited by in corpus (14)
- A Brief Survey of Deep Reinforcement Learning
- An Introduction to Deep Reinforcement Learning
- Learning to Communicate to Solve Riddles with Deep Distributed Recurrent Q-Networks
- Doubly Robust Off-policy Value Evaluation for Reinforcement Learning
- A Deep Recurrent-Reinforcement Learning Method for Intelligent AutoScaling of Serverless Functions
- Neural Predictive Belief Representations
- The Intentional Unintentional Agent: Learning to Solve Many Continuous Control Tasks Simultaneously
- Autonomous Ramp Merge Maneuver Based on Reinforcement Learning with Continuous Action Space
- Causally Correct Partial Models for Reinforcement Learning
- Shaping Belief States with Generative Environment Models for RL
- On the Privacy Risks of Deploying Recurrent Neural Networks in Machine Learning Models
- Robust Actor-Critic Contextual Bandit for Mobile Health (mHealth) Interventions
- Hybrid Supervised Reinforced Model for Dialogue Systems
- How memory architecture affects learning in a simple POMDP: the two-hypothesis testing problem