Deep Recurrent Q-Learning for Partially Observable MDPs
arXiv:1507.06527
Abstract
Deep Reinforcement Learning has yielded proficient controllers for complex tasks. However, these controllers have limited memory and rely on being able to perceive the complete game screen at each decision point. To address these shortcomings, this article investigates the effects of adding recurrency to a Deep Q-Network (DQN) by replacing the first post-convolutional fully-connected layer with a recurrent LSTM. The resulting \textit{Deep Recurrent Q-Network} (DRQN), although capable of seeing only a single frame at each timestep, successfully integrates information through time and replicates DQN's performance on standard Atari games and partially observed equivalents featuring flickering game screens. Additionally, when trained with partial observations and evaluated with incrementally more complete observations, DRQN's performance scales as a function of observability. Conversely, when trained with full observations and evaluated with partial observations, DRQN's performance degrades less than DQN's. Thus, given the same length of history, recurrency is a viable alternative to stacking a history of frames in the DQN's input layer and while recurrency confers no systematic advantage when learning to play the game, the recurrent net can better adapt at evaluation time if the quality of observations changes.
References in corpus (3)
Cited by in corpus (78)
- A Brief Survey of Deep Reinforcement Learning
- Deep Reinforcement Learning for Multi-Agent Systems: A Review of Challenges, Solutions and Applications
- Learning to Communicate with Deep Multi-Agent Reinforcement Learning
- Hierarchical Deep Reinforcement Learning: Integrating Temporal Abstraction and Intrinsic Motivation
- A Survey and Critique of Multiagent Deep Reinforcement Learning
- Predicting Head Movement in Panoramic Video: A Deep Reinforcement Learning Approach
- A Survey of Beam Management for mmWave and THz Communications Towards 6G
- An Application of Deep Reinforcement Learning to Algorithmic Trading
- Federated Reinforcement Learning: Techniques, Applications, and Open Challenges
- Ten Years of Generative Adversarial Nets (GANs): A survey of the state-of-the-art
- Autonomous Unmanned Aerial Vehicle Navigation using Reinforcement Learning: A Systematic Review
- End-to-End Deep Reinforcement Learning for Lane Keeping Assist
- Reinforcement Learning Algorithms: An Overview and Classification
- Opponent Modeling in Deep Reinforcement Learning
- Virtual-to-real Deep Reinforcement Learning: Continuous Control of Mobile Robots for Mapless Navigation
- Guided Deep Reinforcement Learning for Swarm Systems
- ViZDoom Competitions: Playing Doom from Pixels
- Learn to Schedule (LEASCH): A Deep reinforcement learning approach for radio resource scheduling in the 5G MAC layer
- Learning to Communicate to Solve Riddles with Deep Distributed Recurrent Q-Networks
- How to Discount Deep Reinforcement Learning: Towards New Dynamic Strategies
- Towards Cognitive Exploration through Deep Reinforcement Learning for Mobile Robots
- Jointly Learning to Recommend and Advertise
- Deep Reinforcement Learning for Radio Resource Allocation and Management in Next Generation Heterogeneous Wireless Networks: A Survey
- MO-MIX: Multi-Objective Multi-Agent Cooperative Decision-Making With Deep Reinforcement Learning
- Deep Hierarchical Reinforcement Learning Algorithm in Partially Observable Markov Decision Processes
- Recurrent Reinforcement Learning: A Hybrid Approach
- Playing Doom with SLAM-Augmented Deep Reinforcement Learning
- A Deep Reinforcement Learning Framework for Contention-Based Spectrum Sharing
- Value Prediction Network
- A Helmholtz equation solver using unsupervised learning: Application to transcranial ultrasound
- Multi-Agent Reinforcement Learning Based on Representational Communication for Large-Scale Traffic Signal Control
- Deep Multi-agent Reinforcement Learning for Highway On-Ramp Merging in Mixed Traffic
- Recurrent Deterministic Policy Gradient Method for Bipedal Locomotion on Rough Terrain Challenge
- UPDeT: Universal Multi-agent Reinforcement Learning via Policy Decoupling with Transformers
- Machine Learning Methods for Management UAV Flocks -- a Survey
- ALFWorld: Aligning Text and Embodied Environments for Interactive Learning
- WD3: Taming the Estimation Bias in Deep Reinforcement Learning
- Optimal Scheduling in IoT-Driven Smart Isolated Microgrids Based on Deep Reinforcement Learning
- Behavioral decision-making for urban autonomous driving in the presence of pedestrians using Deep Recurrent Q-Network
- General Value Function Networks
- A Deep Recurrent Q Network towards Self-adapting Distributed Microservices architecture
- Aggregating E-commerce Search Results from Heterogeneous Sources via Hierarchical Reinforcement Learning
- Optimized Bacteria are Environmental Prediction Engines
- A Human Mixed Strategy Approach to Deep Reinforcement Learning
- Partially Observable Markov Decision Process for Recommender Systems
- A Survey of Knowledge-based Sequential Decision Making under Uncertainty
- Perceptual Reward Functions
- Traffic Learning and Proactive UAV Trajectory Planning for Data Uplink in Markovian IoT Models
- Experimental Analysis of Reinforcement Learning Techniques for Spectrum Sharing Radar
- Predictive-State Decoders: Encoding the Future into Recurrent Networks
- Learning-to-Ask: Knowledge Acquisition via 20 Questions
- Robust Reinforcement Learning in POMDPs with Incomplete and Noisy Observations
- Why People Skip Music? On Predicting Music Skips using Deep Reinforcement Learning
- Reinforcement Learning for Heterogeneous Teams with PALO Bounds
- Formal Methods for Autonomous Systems
- QR-MIX: Distributional Value Function Factorisation for Cooperative Multi-Agent Reinforcement Learning
- PowerNet: Multi-agent Deep Reinforcement Learning for Scalable Powergrid Control
- Efficient LSTM Training with Eligibility Traces
- Investigating Recurrence and Eligibility Traces in Deep Q-Networks
- Integrating Algorithmic Planning and Deep Learning for Partially Observable Navigation
- Correcting Experience Replay for Multi-Agent Communication
- HARPO: Learning to Subvert Online Behavioral Advertising
- Centralized Model and Exploration Policy for Multi-Agent RL
- Model-Based Episodic Memory Induces Dynamic Hybrid Controls
- Autonomous Curiosity for Real-Time Training Onboard Robotic Agents
- Deep Reinforcement Learning From Raw Pixels in Doom
- A novel control mode of bionic morphing tail based on deep reinforcement learning
- Fighting Copycat Agents in Behavioral Cloning from Observation Histories
- Automatically Learning Fallback Strategies with Model-Free Reinforcement Learning in Safety-Critical Driving Scenarios
- Parallelized Interactive Machine Learning on Autonomous Vehicles
- Easy Monotonic Policy Iteration
- Simultaneous Navigation and Construction Benchmarking Environments
- Augmenting the action space with conventions to improve multi-agent cooperation in Hanabi
- Dynamic Programming for POMDP with Jointly Discrete and Continuous State-Spaces
- Hard Attention Control By Mutual Information Maximization
- Adversarial Reinforcement Learning in Dynamic Channel Access and Power Control
- Towards real-world navigation with deep differentiable planners
- Interpretable UAV Collision Avoidance using Deep Reinforcement Learning