Learning Continuous Control Policies by Stochastic Value Gradients
arXiv:1510.09142
Abstract
We present a unified framework for learning continuous control policies using backpropagation. It supports stochastic control by treating stochasticity in the Bellman equation as a deterministic function of exogenous noise. The product is a spectrum of general policy gradient algorithms that range from model-free methods with value functions to model-based methods without value functions. We use learned models but only require observations from the environment in- stead of observations from model-predicted trajectories, minimizing the impact of compounded model errors. We apply these algorithms first to a toy stochastic control problem and then to several physics-based control problems in simulation. One of these variants, SVG(1), shows the effectiveness of learning models, value functions, and policies simultaneously in continuous domains.
13 pages, NIPS 2015
References in corpus (3)
Cited by in corpus (92)
- Continuous control with deep reinforcement learning
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
- An Introduction to Variational Autoencoders
- Soft Actor-Critic Algorithms and Applications
- High-Dimensional Continuous Control Using Generalized Advantage Estimation
- Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model
- Benchmarking Deep Reinforcement Learning for Continuous Control
- Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems
- Model-Based Reinforcement Learning for Atari
- Continuous Deep Q-Learning with Model-based Acceleration
- Graph networks as learnable physics engines for inference and control
- Learning by Playing - Solving Sparse Reward Tasks from Scratch
- Memory-based control with recurrent neural networks
- Sample Efficient Actor-Critic with Experience Replay
- Differentiable MPC for End-to-end Planning and Control
- Flow: A Modular Learning Framework for Mixed Autonomy Traffic
- Model-Based Value Estimation for Efficient Model-Free Reinforcement Learning
- Dream to Control: Learning Behaviors by Latent Imagination
- DARLA: Improving Zero-Shot Transfer in Reinforcement Learning
- When to Trust Your Model: Model-Based Policy Optimization
- Learning and Transfer of Modulated Locomotor Controllers
- Critic Regularized Regression
- Combining Model-Based and Model-Free Updates for Trajectory-Centric Reinforcement Learning
- Monte Carlo Gradient Estimation in Machine Learning
- Decoupled Neural Interfaces using Synthetic Gradients
- Adaptive dynamic programming for nonaffine nonlinear optimal control problem with state constraints
- Hierarchical visuomotor control of humanoids
- Meta reinforcement learning as task inference
- Freeway Merging in Congested Traffic based on Multipolicy Decision Making with Passive Actor Critic
- Keep Doing What Worked: Behavioral Modelling Priors for Offline Reinforcement Learning
- Learning Implicit Credit Assignment for Cooperative Multi-Agent Reinforcement Learning
- The Mirage of Action-Dependent Baselines in Reinforcement Learning
- Value Prediction Network
- Robust Reinforcement Learning for Continuous Control with Model Misspecification
- Relative Entropy Regularized Policy Iteration
- Constrained Attractor Selection Using Deep Reinforcement Learning
- VIREL: A Variational Inference Framework for Reinforcement Learning
- Maximum Entropy RL (Provably) Solves Some Robust RL Problems
- Constructing Parsimonious Analytic Models for Dynamic Systems via Symbolic Regression
- Learning to Fly via Deep Model-Based Reinforcement Learning
- Recall Traces: Backtracking Models for Efficient Reinforcement Learning
- MBMF: Model-Based Priors for Model-Free Reinforcement Learning
- DeepHealth: Review and challenges of artificial intelligence in health informatics
- Exploiting Hierarchy for Learning and Transfer in KL-regularized RL
- The Wireless Control Plane: An Overview and Directions for Future Research
- On the model-based stochastic value gradient for continuous reinforcement learning
- Boredom-driven curious learning by Homeo-Heterostatic Value Gradients
- Model-aided Deep Reinforcement Learning for Sample-efficient UAV Trajectory Design in IoT Networks
- Trainable Greedy Decoding for Neural Machine Translation
- Provable Benefits of Actor-Critic Methods for Offline Reinforcement Learning
- Learning to Collaborate: Multi-Scenario Ranking via Multi-Agent Reinforcement Learning
- Gradient-Aware Model-based Policy Search
- How to Learn a Useful Critic? Model-based Action-Gradient-Estimator Policy Optimization
- Smoothed Action Value Functions for Learning Gaussian Policies
- Pontryagin Differentiable Programming: An End-to-End Learning and Control Framework
- DyNODE: Neural Ordinary Differential Equations for Dynamics Modeling in Continuous Control
- Actor-critic versus direct policy search: a comparison based on sample complexity
- Hybrid Actor-Critic Reinforcement Learning in Parameterized Action Space
- Interpolated Policy Gradient: Merging On-Policy and Off-Policy Gradient Estimation for Deep Reinforcement Learning
- Deep Reinforcement Learning to Acquire Navigation Skills for Wheel-Legged Robots in Complex Environments
- Counterfactual Credit Assignment in Model-Free Reinforcement Learning
- Active Feature Acquisition with Generative Surrogate Models
- OffWorld Gym: open-access physical robotics environment for real-world reinforcement learning benchmark and research
- Fourier Policy Gradients
- Output-feedback online optimal control for a class of nonlinear systems
- Local Search for Policy Iteration in Continuous Control
- Learning Parametric Closed-Loop Policies for Markov Potential Games
- Deep Model-Based Reinforcement Learning for High-Dimensional Problems, a Survey
- Bridging Imagination and Reality for Model-Based Deep Reinforcement Learning
- Hierarchical Approaches for Reinforcement Learning in Parameterized Action Space
- Model-Based Regularization for Deep Reinforcement Learning with Transcoder Networks
- Only Relevant Information Matters: Filtering Out Noisy Samples to Boost RL
- Generalized Decision Transformer for Offline Hindsight Information Matching
- Iterative Amortized Policy Optimization
- Procedural Generalization by Planning with Self-Supervised World Models
- Expert Level control of Ramp Metering based on Multi-task Deep Reinforcement Learning
- Direct Policy Gradients: Direct Optimization of Policies in Discrete Action Spaces
- Exploring Zero-Shot Emergent Communication in Embodied Multi-Agent Populations
- Fully Distributed Actor-Critic Architecture for Multitask Deep Reinforcement Learning
- Discovering Diverse Athletic Jumping Strategies
- On the Expressivity of Neural Networks for Deep Reinforcement Learning
- Physics-informed Dyna-Style Model-Based Deep Reinforcement Learning for Dynamic Control
- A Unified Off-Policy Evaluation Approach for General Value Function
- Expanding Motor Skills through Relay Neural Networks
- Bayes-Adaptive Deep Model-Based Policy Optimisation
- Aligning Time Series on Incomparable Spaces
- Visual Reaction: Learning to Play Catch with Your Drone
- Generative Temporal Difference Learning for Infinite-Horizon Prediction
- Deterministic Value-Policy Gradients
- Learning How to Solve Bubble Ball
- Model-free Policy Learning with Reward Gradients
- Learning to Reweight Imaginary Transitions for Model-Based Reinforcement Learning