Learning Parameterized Skills
arXiv:1206.6398
Abstract
We introduce a method for constructing skills capable of solving tasks drawn from a distribution of parameterized reinforcement learning problems. The method draws example tasks from a distribution of interest and uses the corresponding learned policies to estimate the topology of the lower-dimensional piecewise-smooth manifold on which the skill policies lie. This manifold models how policy parameters change as task parameters vary. The method identifies the number of charts that compose the manifold and then applies non-linear regression in each chart to construct a parameterized skill by predicting policy parameters from task parameters. We evaluate our method on an underactuated simulated robotic arm tasked with learning to accurately throw darts at a parameterized target location.
Appears in Proceedings of the 29th International Conference on Machine Learning (ICML 2012)
Cited by in corpus (28)
- Hindsight Experience Replay
- One-Shot Visual Imitation Learning via Meta-Learning
- Temporal Difference Models: Model-Free Deep RL for Model-Based Control
- EPOpt: Learning Robust Neural Network Policies Using Model Ensembles
- A Deep Hierarchical Approach to Lifelong Learning in Minecraft
- Zero-Shot Task Generalization with Multi-Task Deep Reinforcement Learning
- Learning to Learn: Meta-Critic Networks for Sample Efficient Learning
- Policy Transfer with Strategy Optimization
- Active choice of teachers, learning strategies and goals for a socially guided intrinsic motivation learner
- Hidden Parameter Markov Decision Processes: A Semiparametric Regression Approach for Discovering Latent Task Parametrizations
- Preparing for the Unknown: Learning a Universal Policy with Online System Identification
- Curiosity-Driven Experience Prioritization via Density Estimation
- Uncertainty Averse Pushing with Model Predictive Path Integral Control
- Hindsight policy gradients
- Energy-Based Hindsight Experience Prioritization
- Distributionally Robust Reinforcement Learning
- Training Agents using Upside-Down Reinforcement Learning
- Deep Imitation Learning for Bimanual Robotic Manipulation
- Fast Adaptation via Policy-Dynamics Value Functions
- SOAC: The Soft Option Actor-Critic Architecture
- Learning Dynamic Robot-to-Human Object Handover from Human Feedback
- Multi-task Learning with Gradient Guided Policy Specialization
- Situational Awareness by Risk-Conscious Skills
- Hierarchical Representation Learning for Markov Decision Processes
- Plan Arithmetic: Compositional Plan Vectors for Multi-Task Control
- Situationally Aware Options
- Policy Search with High-Dimensional Context Variables
- Guided Policy Search for Parameterized Skills using Adverbs