One-Shot Imitation Learning
arXiv:1703.07326
Abstract
Imitation learning has been commonly applied to solve different tasks in isolation. This usually requires either careful feature engineering, or a significant number of samples. This is far from what we desire: ideally, robots should be able to learn from very few demonstrations of any given task, and instantly generalize to new situations of the same task, without requiring task-specific engineering. In this paper, we propose a meta-learning framework for achieving such capability, which we call one-shot imitation learning. Specifically, we consider the setting where there is a very large set of tasks, and each task has many instantiations. For example, a task could be to stack all blocks on a table into a single tower, another task could be to place all blocks on a table into two-block towers, etc. In each case, different instances of the task would consist of different sets of blocks with different initial states. At training time, our algorithm is presented with pairs of demonstrations for a subset of all tasks. A neural net is trained that takes as input one demonstration and the current state (which initially is the initial state of the other demonstration of the pair), and outputs an action with the goal that the resulting sequence of states and actions matches as closely as possible with the second demonstration. At test time, a demonstration of a single instance of a new task is presented, and the neural net is expected to perform well on new instances of this new task. The use of soft attention allows the model to generalize to conditions and tasks unseen in the training data. We anticipate that by training this model on a much greater variety of tasks and settings, we will obtain a general system that can turn any demonstrations into robust policies that can accomplish an overwhelming variety of tasks. Videos available at https://bit.ly/nips2017-oneshot .
References in corpus (6)
- Sequence to Sequence Learning with Neural Networks
- Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation
- Deep Domain Confusion: Maximizing for Domain Invariance
- RL: Fast Reinforcement Learning via Slow Reinforcement Learning
- Learning to reinforcement learn
- Learning Invariant Feature Spaces to Transfer Skills with Reinforcement Learning
Cited by in corpus (40)
- How to Train Your Robot with Deep Reinforcement Learning; Lessons We've Learned
- An Algorithmic Perspective on Imitation Learning
- One-Shot Visual Imitation Learning via Meta-Learning
- Unsupervised Predictive Memory in a Goal-Directed Agent
- Transferring End-to-End Visuomotor Control from Simulation to Real World for a Multi-Stage Task
- One-Shot Imitation from Observing Humans via Domain-Adaptive Meta-Learning
- Mechanical Search: Multi-Step Retrieval of a Target Object Occluded by Clutter
- Reinforcement and Imitation Learning for Diverse Visuomotor Skills
- Relation Networks for Object Detection
- Machine Theory of Mind
- Robust Imitation of Diverse Behaviors
- Time-Contrastive Networks: Self-Supervised Learning from Video
- Few-shot Autoregressive Density Estimation: Towards Learning to Learn Distributions
- Robust Counterfactual Explanations on Graph Neural Networks
- Multi-Modal Imitation Learning from Unstructured Demonstrations using Generative Adversarial Nets
- Data-Efficient Multirobot, Multitask Transfer Learning for Trajectory Tracking
- Virtual-Taobao: Virtualizing Real-world Online Retail Environment for Reinforcement Learning
- Learning Temporal Strategic Relationships using Generative Adversarial Imitation Learning
- Zero-Shot Dialog Generation with Cross-Domain Latent Actions
- Generative One-Shot Learning (GOL): A Semi-Parametric Approach to One-Shot Learning in Autonomous Vision
- Planning on the fast lane: Learning to interact using attention mechanisms in path integral inverse reinforcement learning
- A Survey of Behavior Learning Applications in Robotics -- State of the Art and Perspectives
- Meta Inverse Reinforcement Learning via Maximum Reward Sharing for Human Motion Analysis
- Cascade Attribute Learning Network
- Learning Deep Parameterized Skills from Demonstration for Re-targetable Visuomotor Control
- CrowdTransfer: Enabling Crowd Knowledge Transfer in AIoT Community
- One-Shot Learning on Attributed Sequences
- Learning with Stochastic Guidance for Navigation
- Scene learning, recognition and similarity detection in a fuzzy ontology via human examples
- Hybrid Reinforcement Learning with Expert State Sequences
- MimicBot: Combining Imitation and Reinforcement Learning to win in Bot Bowl
- Learning to Play by Imitating Humans
- Guided Exploration with Proximal Policy Optimization using a Single Demonstration
- AGENT: A Benchmark for Core Psychological Reasoning
- Stable Object Reorientation using Contact Plane Registration
- Learning to Imagine Manipulation Goals for Robot Task Planning
- Dynamic Regret Convergence Analysis and an Adaptive Regularization Algorithm for On-Policy Robot Imitation Learning
- General AI Challenge - Round One: Gradual Learning
- Bottom-Up Skill Discovery from Unsegmented Demonstrations for Long-Horizon Robot Manipulation
- Learning with Training Wheels: Speeding up Training with a Simple Controller for Deep Reinforcement Learning