O2A: One-shot Observational learning with Action vectors
arXiv:1810.07483 · doi:10.3389/frobt.2021.686368
Abstract
We present O2A, a novel method for learning to perform robotic manipulation tasks from a single (one-shot) third-person demonstration video. To our knowledge, it is the first time this has been done for a single demonstration. The key novelty lies in pre-training a feature extractor for creating a perceptual representation for actions that we call 'action vectors'. The action vectors are extracted using a 3D-CNN model pre-trained as an action classifier on a generic action dataset. The distance between the action vectors from the observed third-person demonstration and trial robot executions is used as a reward for reinforcement learning of the demonstrated task. We report on experiments in simulation and on a real robot, with changes in viewpoint of observation, properties of the objects involved, scene background and morphology of the manipulator between the demonstration and the learning domains. O2A outperforms baseline approaches under different domain shifts and has comparable performance with an oracle (that uses an ideal reward function).
References in corpus (12)
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- DeepMimic: Example-Guided Deep Reinforcement Learning of Physics-Based Character Skills
- Deep Reinforcement Learning that Matters
- Third-Person Imitation Learning
- Feature Representation in Convolutional Neural Networks
- One-Shot Hierarchical Imitation Learning of Compound Visuomotor Tasks
- Multiple Interactions Made Easy (MIME): Large Scale Demonstrations Data for Imitation
- Third-Person Visual Imitation Learning via Decoupled Hierarchical Controller
- Graph-Structured Visual Imitation
- Vision-based Robot Manipulation Learning via Human Demonstrations
- Pushing Fast and Slow: Task-Adaptive Planning for Non-prehensile Manipulation Under Uncertainty
- Defining the problem of Observation Learning