Self-supervised Reinforcement Learning with Independently Controllable Subgoals
arXiv:2109.04150
Abstract
To successfully tackle challenging manipulation tasks, autonomous agents must learn a diverse set of skills and how to combine them. Recently, self-supervised agents that set their own abstract goals by exploiting the discovered structure in the environment were shown to perform well on many different tasks. In particular, some of them were applied to learn basic manipulation skills in compositional multi-object environments. However, these methods learn skills without taking the dependencies between objects into account. Thus, the learned skills are difficult to combine in realistic environments. We propose a novel self-supervised agent that estimates relations between environment components and uses them to independently control different parts of the environment state. In addition, the estimated relations between objects can be used to decompose a complex goal into a compatible sequence of subgoals. We show that, by using this framework, an agent can efficiently and automatically learn manipulation tasks in multi-object environments with different relations between objects.
References in corpus (12)
- Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling
- Interaction Networks for Learning about Objects, Relations and Physics
- Active Learning of Inverse Models with Intrinsically Motivated Goal Exploration in Robots
- FeUdal Networks for Hierarchical Reinforcement Learning
- MONet: Unsupervised Scene Decomposition and Representation
- Amortized Causal Discovery: Learning to Infer Causal Graphs from Time-Series Data
- Causal Discovery in Physical Systems from Videos
- SCALOR: Generative World Models with Scalable Object Representations
- Grounding Language to Autonomously-Acquired Skills via Goal Generation
- Hierarchical Policy Learning is Sensitive to Goal Space Design
- Causal Influence Detection for Improving Efficiency in Reinforcement Learning
- ROLL: Visual Self-Supervised Reinforcement Learning with Object Reasoning