Unsupervised Control Through Non-Parametric Discriminative Rewards
arXiv:1811.11359
Abstract
Learning to control an environment without hand-crafted rewards or expert data remains challenging and is at the frontier of reinforcement learning research. We present an unsupervised learning algorithm to train agents to achieve perceptually-specified goals using only a stream of observations and actions. Our agent simultaneously learns a goal-conditioned policy and a goal achievement reward function that measures how similar a state is to the goal state. This dual optimization leads to a co-operative game, giving rise to a learned reward function that reflects similarity in controllable aspects of the environment instead of distance in the space of observations. We demonstrate the efficacy of our agent to learn, in an unsupervised manner, to reach a diverse set of goals on three domains -- Atari, the DeepMind Control Suite and DeepMind Lab.
10 pages + references & 5 page appendix
Cited by in corpus (24)
- A survey on intrinsic motivation in reinforcement learning
- RODE: Learning Roles to Decompose Multi-Agent Tasks
- An information-theoretic perspective on intrinsic motivation in reinforcement learning: a survey
- The Free Energy Principle for Perception and Action: A Deep Learning Perspective
- Hierarchical Reinforcement Learning By Discovering Intrinsic Options
- Contextual Imagined Goals for Self-Supervised Robotic Learning
- Goal-directed Planning and Goal Understanding by Active Inference: Evaluation Through Simulated and Physical Robot Experiments
- Explore, Discover and Learn: Unsupervised Discovery of State-Covering Skills
- Hyperparameter Auto-tuning in Self-Supervised Robotic Learning
- Behavior From the Void: Unsupervised Active Pre-Training
- From Scratch to Sketch: Deep Decoupled Hierarchical Reinforcement Learning for Robotic Sketching Agent
- Solving Compositional Reinforcement Learning Problems via Task Reduction
- Learning Efficient Representation for Intrinsic Motivation
- GRIMGEP: Learning Progress for Robust Goal Sampling in Visual Deep Reinforcement Learning
- Learning more skills through optimistic exploration
- Planning from Pixels using Inverse Dynamics Models
- Variational Empowerment as Representation Learning for Goal-Based Reinforcement Learning
- Self-supervised Visual Reinforcement Learning with Object-centric Representations
- Mutual Information State Intrinsic Control
- Policy Information Capacity: Information-Theoretic Measure for Task Complexity in Deep Reinforcement Learning
- The Information Geometry of Unsupervised Reinforcement Learning
- Contrastive Active Inference
- Adaptive Multi-Goal Exploration
- Direct then Diffuse: Incremental Unsupervised Skill Discovery for State Covering and Goal Reaching