Hierarchical Foresight: Self-Supervised Learning of Long-Horizon Tasks via Visual Subgoal Generation
arXiv:1909.05829
Abstract
Video prediction models combined with planning algorithms have shown promise in enabling robots to learn to perform many vision-based tasks through only self-supervision, reaching novel goals in cluttered scenes with unseen objects. However, due to the compounding uncertainty in long horizon video prediction and poor scalability of sampling-based planning optimizers, one significant limitation of these approaches is the ability to plan over long horizons to reach distant goals. To that end, we propose a framework for subgoal generation and planning, hierarchical visual foresight (HVF), which generates subgoal images conditioned on a goal image, and uses them for planning. The subgoal images are directly optimized to decompose the task into easy to plan segments, and as a result, we observe that the method naturally identifies semantically meaningful states as subgoals. Across three out of four simulated vision-based manipulation tasks, we find that our method achieves nearly a 200% performance improvement over planning without subgoals and model-free RL approaches. Further, our experiments illustrate that our approach extends to real, cluttered visual scenes. Project page: https://sites.google.com/stanford.edu/hvf
16 pages, 9 figures
References in corpus (14)
- QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation
- Hindsight Experience Replay
- Visual Reinforcement Learning with Imagined Goals
- Visual Foresight: Model-Based Deep Reinforcement Learning for Vision-Based Robotic Control
- Transferring End-to-End Visuomotor Control from Simulation to Real World for a Multi-Stage Task
- Self-Supervised Visual Planning with Temporal Skip Connections
- Sim-to-Real Reinforcement Learning for Deformable Object Manipulation
- Stochastic Variational Video Prediction
- Hierarchical Reinforcement Learning with Hindsight
- Learning Latent Plans from Play
- Deep Visual Foresight for Planning Robot Motion
- Neural Task Programming: Learning to Generalize Across Hierarchical Tasks
- Time Reversal as Self-Supervision
- Robot Motion Planning in Learned Latent Spaces
Cited by in corpus (5)
- COG: Connecting New Skills to Past Experience with Offline Reinforcement Learning
- Long-Horizon Visual Planning with Goal-Conditioned Hierarchical Predictors
- Disentangling causal effects for hierarchical reinforcement learning
- Learning Compositional Neural Programs for Continuous Control
- Learning a Skill-sequence-dependent Policy for Long-horizon Manipulation Tasks