A Brief Survey of Deep Reinforcement Learning
arXiv:1708.05866 · doi:10.1109/MSP.2017.2743240
Abstract
Deep reinforcement learning is poised to revolutionise the field of AI and represents a step towards building autonomous systems with a higher level understanding of the visual world. Currently, deep learning is enabling reinforcement learning to scale to problems that were previously intractable, such as learning to play video games directly from pixels. Deep reinforcement learning algorithms are also applied to robotics, allowing control policies for robots to be learned directly from camera inputs in the real world. In this survey, we begin with an introduction to the general field of reinforcement learning, then progress to the main streams of value-based and policy-based methods. Our survey will cover central algorithms in deep reinforcement learning, including the deep -network, trust region policy optimisation, and asynchronous advantage actor-critic. In parallel, we highlight the unique advantages of deep neural networks, focusing on visual understanding via reinforcement learning. To conclude, we describe several current areas of research within the field.
IEEE Signal Processing Magazine, Special Issue on Deep Learning for Image Understanding (arXiv extended version)
References in corpus (19)
- Distilling the Knowledge in a Neural Network
- Neural Architecture Search with Reinforcement Learning
- StarCraft II: A New Challenge for Reinforcement Learning
- Emergence of Locomotion Behaviours in Rich Environments
- RL: Fast Reinforcement Learning via Slow Reinforcement Learning
- Massively Parallel Methods for Deep Reinforcement Learning
- Learning to Navigate in Complex Environments
- Optimizing Dialogue Management with Reinforcement Learning: Experiments with the NJFun System
- Multi-agent Reinforcement Learning in Sequential Social Dilemmas
- Reinforcement Learning with Unsupervised Auxiliary Tasks
- FeUdal Networks for Hierarchical Reinforcement Learning
- Distral: Robust Multitask Reinforcement Learning
- Schema Networks: Zero-shot Transfer with a Generative Causal Model of Intuitive Physics
- Towards Deep Symbolic Reinforcement Learning
- Episodic Exploration for Deep Deterministic Policies: An Application to StarCraft Micromanagement Tasks
- TorchCraft: a Library for Machine Learning Research on Real-Time Strategy Games
- Learning model-based planning from scratch
- DeepMind Lab
- Model-based Adversarial Imitation Learning
Cited by in corpus (74)
- Exploration in Deep Reinforcement Learning: A Survey
- Deep Learning Techniques for Future Intelligent Cross-Media Retrieval
- Learning Combinatorial Optimization on Graphs: A Survey with Applications to Networking
- AI-enabled Future Wireless Networks: Challenges, Opportunities and Open Issues
- A Survey of Deep Reinforcement Learning in Video Games
- Deep Reinforcement Learning Control for Radar Detection and Tracking in Congested Spectral Environments
- A Review on Computational Intelligence Techniques in Cloud and Edge Computing
- Intelligent problem-solving as integrated hierarchical reinforcement learning
- Machine and Deep Learning for IoT Security and Privacy: Applications, Challenges, and Future Directions
- Multi-UAV Path Learning for Age and Power Optimization in IoT with UAV Battery Recharge
- Empowering Things with Intelligence: A Survey of the Progress, Challenges, and Opportunities in Artificial Intelligence of Things
- A Comprehensive Survey of Incentive Mechanism for Federated Learning
- From Chess and Atari to StarCraft and Beyond: How Game AI is Driving the World of AI
- Neurosymbolic Reinforcement Learning and Planning: A Survey
- Avoiding Catastrophe: Active Dendrites Enable Multi-Task Learning in Dynamic Environments
- Adaptive Control of Resource Flow to Optimize Construction Work and Cash Flow via Online Deep Reinforcement Learning
- Machine learning \& artificial intelligence in the quantum domain
- Tactile based Intelligence Touch Technology in IoT configured WCN in B5G/6G-A Survey
- Trading-off Accuracy and Energy of Deep Inference on Embedded Systems: A Co-Design Approach
- RLAS-BIABC: A Reinforcement Learning-Based Answer Selection Using the BERT Model Boosted by an Improved ABC Algorithm
- A Comprehensive Survey of Machine Learning Applied to Radar Signal Processing
- Optimal active particle navigation meets machine learning
- Learning to Selectively Transfer: Reinforced Transfer Learning for Deep Text Matching
- A Survey of Deep Reinforcement Learning in Recommender Systems: A Systematic Review and Future Directions
- A Learning-Based Trajectory Planning of Multiple UAVs for AoI Minimization in IoT Networks
- Deep Learning for Insider Threat Detection: Review, Challenges and Opportunities
- Balancing a CartPole System with Reinforcement Learning -- A Tutorial
- Specialization in Hierarchical Learning Systems
- A Review of Meta-Reinforcement Learning for Deep Neural Networks Architecture Search
- Bayesian Optimization Meets Riemannian Manifolds in Robot Learning
- Deep Reinforcement Learning Based High-level Driving Behavior Decision-making Model in Heterogeneous Traffic
- Explainable Artificial Intelligence (XAI) for 6G: Improving Trust between Human and Machine
- Is Deep Reinforcement Learning Ready for Practical Applications in Healthcare? A Sensitivity Analysis of Duel-DDQN for Hemodynamic Management in Sepsis Patients
- Transformer Based Reinforcement Learning For Games
- Deep Actor-Critic Learning for Distributed Power Control in Wireless Mobile Networks
- A Reinforcement Learning Approach for an IRS-assisted NOMA Network
- Increasing performance of electric vehicles in ride-hailing services using deep reinforcement learning
- Adaptive Road Configurations for Improved Autonomous Vehicle-Pedestrian Interactions using Reinforcement Learning
- HyperNCA: Growing Developmental Networks with Neural Cellular Automata
- Safer Deep RL with Shallow MCTS: A Case Study in Pommerman
- Zeroth-order Deterministic Policy Gradient
- Computational Rational Engineering and Development: Synergies and Opportunities
- Green Deep Reinforcement Learning for Radio Resource Management: Architecture, Algorithm Compression and Challenge
- Independent Natural Policy Gradient Always Converges in Markov Potential Games
- Are Gradient-based Saliency Maps Useful in Deep Reinforcement Learning?
- Diversity-based Trajectory and Goal Selection with Hindsight Experience Replay
- Reinforcement Learning for Intelligent Healthcare Systems: A Comprehensive Survey
- Explainable Deep Reinforcement Learning Using Introspection in a Non-episodic Task
- Intelligent Network Slicing for V2X Services Towards 5G
- Reinforcing Medical Image Classifier to Improve Generalization on Small Datasets
- Active Screening for Recurrent Diseases: A Reinforcement Learning Approach
- MimicBot: Combining Imitation and Reinforcement Learning to win in Bot Bowl
- Learning Knowledge Graph-based World Models of Textual Environments
- Reinforcement Learning for a Cellular Internet of UAVs: Protocol Design, Trajectory Control, and Resource Management
- Sample-Efficient Reinforcement Learning with Maximum Entropy Mellowmax Episodic Control
- Approximation of Optimal Control Surfaces for Skew-Symmetric Evolutionary Game Dynamics
- Convolutional Reservoir Computing for World Models
- Reinforcement Learning with Convolutional Reservoir Computing
- A New Approach for Tactical Decision Making in Lane Changing: Sample Efficient Deep Q Learning with a Safety Feedback Reward
- Model-Free Design of Stochastic LQR Controller from Reinforcement Learning and Primal-Dual Optimization Perspective
- Modeling Worlds in Text
- Episodic Self-Imitation Learning with Hindsight
- Interpretable Option Discovery using Deep Q-Learning and Variational Autoencoders
- Interpretable UAV Collision Avoidance using Deep Reinforcement Learning
- Decentralized Deep Reinforcement Learning for Network Level Traffic Signal Control
- Reinforcement Learning Architectures: SAC, TAC, and ESAC
- On Deep Learning for Radio Resource Management in A Non-stationary Radio Environment
- Communication-Computation Efficient Device-Edge Co-Inference via AutoML
- SeaPearl: A Constraint Programming Solver guided by Reinforcement Learning
- Techniques Toward Optimizing Viewability in RTB Ad Campaigns Using Reinforcement Learning
- RLCache: Automated Cache Management Using Reinforcement Learning
- Generalized Operating Procedure for Deep Learning: an Unconstrained Optimal Design Perspective
- Partially Observable Planning and Learning for Systems with Non-Uniform Dynamics
- Continuous Deep Q-Learning with Simulator for Stabilization of Uncertain Discrete-Time Systems