Continuous Deep Q-Learning with Model-based Acceleration
arXiv:1603.00748
Abstract
Model-free reinforcement learning has been successfully applied to a range of challenging problems, and has recently been extended to handle large neural network policies and value functions. However, the sample complexity of model-free algorithms, particularly when using high-dimensional function approximators, tends to limit their applicability to physical systems. In this paper, we explore algorithms and representations to reduce the sample complexity of deep reinforcement learning for continuous control tasks. We propose two complementary techniques for improving the efficiency of such algorithms. First, we derive a continuous variant of the Q-learning algorithm, which we call normalized adantage functions (NAF), as an alternative to the more commonly used policy gradient and actor-critic methods. NAF representation allows us to apply Q-learning with experience replay to continuous tasks, and substantially improves performance on a set of simulated robotic control tasks. To further improve the efficiency of our approach, we explore the use of learned models for accelerating model-free reinforcement learning. We show that iteratively refitted local linear models are especially effective for this, and demonstrate substantially faster learning on domains where such models are applicable.
References in corpus (1)
Cited by in corpus (128)
- A Brief Survey of Deep Reinforcement Learning
- A Survey of Deep Learning Techniques for Autonomous Driving
- Deep Reinforcement Learning for Multi-Agent Systems: A Review of Challenges, Solutions and Applications
- How to Train Your Robot with Deep Reinforcement Learning; Lessons We've Learned
- Deep Reinforcement Learning: An Overview
- Reinforcement Learning with Deep Energy-Based Policies
- A Survey of Deep RL and IL for Autonomous Driving Policy Learning
- Robust active flow control over a range of Reynolds numbers using an artificial neural network trained through deep reinforcement learning
- Differentiable MPC for End-to-end Planning and Control
- Learning Particle Dynamics for Manipulating Rigid Bodies, Deformable Objects, and Fluids
- Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning
- Model-Based Value Estimation for Efficient Model-Free Reinforcement Learning
- Algorithmic Framework for Model-based Deep Reinforcement Learning with Theoretical Guarantees
- When to Trust Your Model: Model-Based Policy Optimization
- A Generalized Algorithm for Multi-Objective Reinforcement Learning and Policy Adaptation
- FACMAC: Factored Multi-Agent Centralised Policy Gradients
- Learning to Adapt in Dynamic, Real-World Environments Through Meta-Reinforcement Learning
- Q-Prop: Sample-Efficient Policy Gradient with An Off-Policy Critic
- A Review of Robot Learning for Manipulation: Challenges, Representations, and Algorithms
- Distributed Reinforcement Learning for Privacy-Preserving Dynamic Edge Caching
- Exploring Model-based Planning with Policy Networks
- ChainerRL: A Deep Reinforcement Learning Library
- Falsification of Cyber-Physical Systems Using Deep Reinforcement Learning
- Stochastic Latent Actor-Critic: Deep Reinforcement Learning with a Latent Variable Model
- Learning Stable Deep Dynamics Models
- Residual Policy Learning
- rlpyt: A Research Code Base for Deep Reinforcement Learning in PyTorch
- Deployment-Efficient Reinforcement Learning via Model-Based Offline Optimization
- Model-Based Reinforcement Learning with Adversarial Training for Online Recommendation
- Responsive Safety in Reinforcement Learning by PID Lagrangian Methods
- The Mirage of Action-Dependent Baselines in Reinforcement Learning
- Boosting Soft Actor-Critic: Emphasizing Recent Experience without Forgetting the Past
- Robust Deep Reinforcement Learning against Adversarial Perturbations on State Observations
- A Deep Learning Approach for Joint Video Frame and Reward Prediction in Atari Games
- Cryptocurrency Trading: A Comprehensive Survey
- Q-Learning in enormous action spaces via amortized approximate maximization
- A Tutorial on Ultra-Reliable and Low-Latency Communications in 6G: Integrating Domain Knowledge into Deep Learning
- Maximum Entropy RL (Provably) Solves Some Robust RL Problems
- Constructing Parsimonious Analytic Models for Dynamic Systems via Symbolic Regression
- The Effect of Planning Shape on Dyna-style Planning in High-dimensional State Spaces
- Context-aware Dynamics Model for Generalization in Model-Based Reinforcement Learning
- Recall Traces: Backtracking Models for Efficient Reinforcement Learning
- ARCHER: Aggressive Rewards to Counter bias in Hindsight Experience Replay
- Tsallis Reinforcement Learning: A Unified Framework for Maximum Entropy Reinforcement Learning
- Deep Reinforcement Learning and Transportation Research: A Comprehensive Review
- On the model-based stochastic value gradient for continuous reinforcement learning
- Combining Q-Learning and Search with Amortized Value Estimates
- MBRL-Lib: A Modular Library for Model-based Reinforcement Learning
- Reinforcement Learning with Perturbed Rewards
- Watch, Try, Learn: Meta-Learning from Demonstrations and Reward
- Contextual Imagined Goals for Self-Supervised Robotic Learning
- Hamilton-Jacobi Deep Q-Learning for Deterministic Continuous-Time Systems with Lipschitz Continuous Controls
- Remember and Forget for Experience Replay
- Striving for Simplicity and Performance in Off-Policy DRL: Output Normalization and Non-Uniform Sampling
- CAQL: Continuous Action Q-Learning
- An initial attempt of combining visual selective attention with deep reinforcement learning
- LADDER: A Human-Level Bidding Agent for Large-Scale Real-Time Online Auctions
- Automated Driving Maneuvers under Interactive Environment based on Deep Reinforcement Learning
- Challenges of Applying Deep Reinforcement Learning in Dynamic Dispatching
- Case Study: Verifying the Safety of an Autonomous Racing Car with a Neural Network Controller
- Propagation Networks for Model-Based Control Under Partial Observation
- Pontryagin Differentiable Programming: An End-to-End Learning and Control Framework
- Model-based Lookahead Reinforcement Learning
- Augmented World Models Facilitate Zero-Shot Dynamics Generalization From a Single Offline Environment
- Quadratic Q-network for Learning Continuous Control for Autonomous Vehicles
- Using RGB Image as Visual Input for Mapless Robot Navigation
- Continuous-action Reinforcement Learning for Playing Racing Games: Comparing SPG to PPO
- Active Feature Acquisition with Generative Surrogate Models
- OffWorld Gym: open-access physical robotics environment for real-world reinforcement learning benchmark and research
- EMaQ: Expected-Max Q-Learning Operator for Simple Yet Effective Offline and Online RL
- Deep Reinforcement Learning with Embedded LQR Controllers
- Deep Residual Reinforcement Learning
- Imagined Value Gradients: Model-Based Policy Optimization with Transferable Latent Dynamics Models
- Uncertainty-sensitive Learning and Planning with Ensembles
- Efficient and Robust Reinforcement Learning with Uncertainty-based Value Expansion
- Off-Policy Actor-Critic in an Ensemble: Achieving Maximum General Entropy and Effective Environment Exploration in Deep Reinforcement Learning
- Bach2Bach: Generating Music Using A Deep Reinforcement Learning Approach
- Generalized Decision Transformer for Offline Hindsight Information Matching
- Steady State Analysis of Episodic Reinforcement Learning
- Few Shot System Identification for Reinforcement Learning
- Tutoring Reinforcement Learning via Feedback Control
- Quinoa: a Q-function You Infer Normalized Over Actions
- Human-in-the-Loop Methods for Data-Driven and Reinforcement Learning Systems
- MobILE: Model-Based Imitation Learning From Observation Alone
- Policy Search by Target Distribution Learning for Continuous Control
- Competitive Experience Replay
- Continuous Control for Searching and Planning with a Learned Model
- Target-Based Temporal Difference Learning
- Control as Hybrid Inference
- AutoEG: Automated Experience Grafting for Off-Policy Deep Reinforcement Learning
- Learning Value Functions in Deep Policy Gradients using Residual Variance
- Control-Tutored Reinforcement Learning
- Towards Generalization and Data Efficient Learning of Deep Robotic Grasping
- Causal World Models by Unsupervised Deconfounding of Physical Dynamics
- Policy Information Capacity: Information-Theoretic Measure for Task Complexity in Deep Reinforcement Learning
- Deep Reinforced Attention Learning for Quality-Aware Visual Recognition
- Continuous Control with Action Quantization from Demonstrations
- Imaginary Hindsight Experience Replay: Curious Model-based Learning for Sparse Reward Tasks
- Fuzzy Tiling Activations: A Simple Approach to Learning Sparse Representations Online
- Fixed Points in Cyber Space: Rethinking Optimal Evasion Attacks in the Age of AI-NIDS
- Personalized Cancer Chemotherapy Schedule: a numerical comparison of performance and robustness in model-based and model-free scheduling methodologies
- A Deep Reinforcement Learning Architecture for Multi-stage Optimal Control
- Frequency-based Search-control in Dyna
- Object-Oriented Dynamics Learning through Multi-Level Abstraction
- Importance Sampling based Exploration in Q Learning
- Implicitly Regularized RL with Implicit Q-Values
- Control-Tutored Reinforcement Learning: an application to the Herding Problem
- Lachesis: Automatic Partitioning for UDF-Centric Analytics
- Extended Radial Basis Function Controller for Reinforcement Learning
- Iterative Model-Based Reinforcement Learning Using Simulations in the Differentiable Neural Computer
- Costate-focused models for reinforcement learning
- Critic PI2: Master Continuous Planning via Policy Improvement with Path Integrals and Deep Actor-Critic Reinforcement Learning
- Deep Reinforcement Learning for Personalized Search Story Recommendation
- Efficient Reinforcement Learning Development with RLzoo
- A2: Extracting Cyclic Switchings from DOB-nets for Rejecting Excessive Disturbances
- Learning to Reach, Swim, Walk and Fly in One Trial: Data-Driven Control with Scarce Data and Side Information
- Lineage Evolution Reinforcement Learning
- Planning with Expectation Models for Control
- Learning with Training Wheels: Speeding up Training with a Simple Controller for Deep Reinforcement Learning
- A Contraction Approach to Model-based Reinforcement Learning
- Eden: A Unified Environment Framework for Booming Reinforcement Learning Algorithms
- Joint Perception and Control as Inference with an Object-based Implementation
- Co-Adaptation of Algorithmic and Implementational Innovations in Inference-based Deep Reinforcement Learning
- Deterministic Value-Policy Gradients
- Multi-Agent Cooperative Bidding Games for Multi-Objective Optimization in e-Commercial Sponsored Search
- Training an Interactive Helper
- Decentralized Deep Reinforcement Learning for Network Level Traffic Signal Control
- Visual Reaction: Learning to Play Catch with Your Drone