Cross-Domain Perceptual Reward Functions
arXiv:1705.09045
Abstract
In reinforcement learning, we often define goals by specifying rewards within desirable states. One problem with this approach is that we typically need to redefine the rewards each time the goal changes, which often requires some understanding of the solution in the agents environment. When humans are learning to complete tasks, we regularly utilize alternative sources that guide our understanding of the problem. Such task representations allow one to specify goals on their own terms, thus providing specifications that can be appropriately interpreted across various environments. This motivates our own work, in which we represent goals in environments that are different from the agents. We introduce Cross-Domain Perceptual Reward (CDPR) functions, learned rewards that represent the visual similarity between an agents state and a cross-domain goal image. We report results for learning the CDPRs with a deep neural network and using them to solve two tasks with deep reinforcement learning.
A shorter version of this paper was accepted to RLDM (http://rldm.org/rldm2017/)
References in corpus (7)
- Neural Architecture Search with Reinforcement Learning
- Domain Separation Networks
- Domain Randomization for Transferring Deep Neural Networks from Simulation to the Real World
- Target-driven Visual Navigation in Indoor Scenes using Deep Reinforcement Learning
- Third-Person Imitation Learning
- Generalizing Skills with Semi-Supervised Reinforcement Learning
- Beating Atari with Natural Language Guided Reinforcement Learning