Third-Person Imitation Learning
arXiv:1703.01703
Abstract
Reinforcement learning (RL) makes it possible to train agents capable of achieving sophisticated goals in complex and uncertain environments. A key difficulty in reinforcement learning is specifying a reward function for the agent to optimize. Traditionally, imitation learning in RL has been used to overcome this problem. Unfortunately, hitherto imitation learning methods tend to require that demonstrations are supplied in the first-person: the agent is provided with a sequence of states and a specification of the actions that it should have taken. While powerful, this kind of imitation learning is limited by the relatively hard problem of collecting first-person demonstrations. Humans address this problem by learning from third-person demonstrations: they observe other humans perform tasks, infer the task, and accomplish the same task themselves. In this paper, we present a method for unsupervised third-person imitation learning. Here third-person refers to training an agent to correctly achieve a simple goal in a simple environment when it is provided a demonstration of a teacher achieving the same goal but from a different viewpoint; and unsupervised refers to the fact that the agent receives only these third-person demonstrations, and is not provided a correspondence between teacher states and student states. Our methods primary insight is that recent advances from domain confusion can be utilized to yield domain agnostic features which are crucial during the training process. To validate our approach, we report successful experiments on learning from third-person demonstrations in a pointmass domain, a reacher domain, and inverted pendulum.
Only changed the abstract to remove unneeded hyphens
References in corpus (2)
Cited by in corpus (55)
- A Brief Survey of Deep Reinforcement Learning
- Deep Reinforcement Learning: An Overview
- An Algorithmic Perspective on Imitation Learning
- InfoGAIL: Interpretable Imitation Learning from Visual Demonstrations
- Transferring End-to-End Visuomotor Control from Simulation to Real World for a Multi-Stage Task
- One-Shot Imitation from Observing Humans via Domain-Adaptive Meta-Learning
- A Review of Robot Learning for Manipulation: Challenges, Representations, and Algorithms
- Aligning Superhuman AI with Human Behavior: Chess as a Model System
- A Survey of Deep Network Solutions for Learning Control in Robotics: From Reinforcement to Imitation
- Robust Imitation of Diverse Behaviors
- Imitating Latent Policies from Observation
- Deep Imitation Learning for Complex Manipulation Tasks from Virtual Reality Teleoperation
- AI-GAs: AI-generating algorithms, an alternate paradigm for producing general artificial intelligence
- Time-Contrastive Networks: Self-Supervised Learning from Video
- Imitating Interactive Intelligence
- Imitation from Observation: Learning to Imitate Behaviors from Raw Video via Context Translation
- Goal-conditioned Imitation Learning
- Virtual-Taobao: Virtualizing Real-world Online Retail Environment for Reinforcement Learning
- Mutual Alignment Transfer Learning
- Learning Temporal Strategic Relationships using Generative Adversarial Imitation Learning
- Graph-Structured Visual Imitation
- Deep Reinforcement Learning and Transportation Research: A Comprehensive Review
- Learning Norms from Stories: A Prior for Value Aligned Agents
- Imitation Learning: Progress, Taxonomies and Challenges
- XIRL: Cross-embodiment Inverse Reinforcement Learning
- State-only Imitation with Transition Dynamics Mismatch
- All by Myself: Learning Individualized Competitive Behaviour with a Contrastive Reinforcement Learning optimization
- Cross-Domain Perceptual Reward Functions
- Adversarial Imitation Learning from Incomplete Demonstrations
- O2A: One-shot Observational learning with Action vectors
- Zero-shot Imitation Learning from Demonstrations for Legged Robot Visual Navigation
- ADAIL: Adaptive Adversarial Imitation Learning
- Reinforcement Learning with Videos: Combining Offline Observations with Interaction
- Leveraging Human Guidance for Deep Reinforcement Learning Tasks
- Human-guided Robot Behavior Learning: A GAN-assisted Preference-based Reinforcement Learning Approach
- Reinforced Imitation in Heterogeneous Action Space
- Domain-Robust Visual Imitation Learning with Mutual Information Constraints
- Interactive Language Acquisition with One-shot Visual Concept Learning through a Conversational Game
- Defining the problem of Observation Learning
- Hybrid Reinforcement Learning with Expert State Sequences
- Hierarchically Decoupled Imitation for Morphological Transfer
- Cross-Domain Imitation Learning with a Dual Structure
- Interactive Learning from Activity Description
- Domain-Adversarial and Conditional State Space Model for Imitation Learning
- Cross-Domain Imitation Learning via Optimal Transport
- Provably Efficient Third-Person Imitation from Offline Observation
- Towards Empathic Deep Q-Learning
- Learning Cloth Folding Tasks with Refined Flow Based Spatio-Temporal Graphs
- PLOTS: Procedure Learning from Observations using Subtask Structure
- Task Transfer by Preference-Based Cost Learning
- Manipulator-Independent Representations for Visual Imitation
- Motion Reasoning for Goal-Based Imitation Learning
- A Bayesian Approach to Identifying Representational Errors
- The Goofus & Gallant Story Corpus for Practical Value Alignment
- HILONet: Hierarchical Imitation Learning from Non-Aligned Observations