Unsupervised Video Object Segmentation for Deep Reinforcement Learning
arXiv:1805.07780
Abstract
We present a new technique for deep reinforcement learning that automatically detects moving objects and uses the relevant information for action selection. The detection of moving objects is done in an unsupervised way by exploiting structure from motion. Instead of directly learning a policy from raw images, the agent first learns to detect and segment moving objects by exploiting flow information in video sequences. The learned representation is then used to focus the policy of the agent on the moving objects. Over time, the agent identifies which objects are critical for decision making and gradually builds a policy based on relevant moving objects. This approach, which we call Motion-Oriented REinforcement Learning (MOREL), is demonstrated on a suite of Atari games where the ability to detect moving objects reduces the amount of interaction needed with the environment to obtain a good policy. Furthermore, the resulting policy is more interpretable than policies that directly map images to actions or values with a black box neural network. We can gain insight into the policy by inspecting the segmentation and motion of each object detected by the agent. This allows practitioners to confirm whether a policy is making decisions based on sensible information.
References in corpus (8)
- Playing Atari with Deep Reinforcement Learning
- Image Segmentation in Video Sequences: A Probabilistic Approach
- Scalable trust-region method for deep reinforcement learning using Kronecker-factored approximation
- Learning to Navigate in Complex Environments
- SfM-Net: Learning of Structure and Motion from Video
- Unsupervised Learning of Depth and Ego-Motion from Video
- Schema Networks: Zero-shot Transfer with a Generative Causal Model of Intuitive Physics
- Learning Multimodal Transition Dynamics for Model-Based Reinforcement Learning
Cited by in corpus (9)
- Unsupervised Learning of Object Keypoints for Perception and Control
- Local and Global Explanations of Agent Behavior: Integrating Strategy Summaries with Saliency Maps
- Entity Abstraction in Visual Model-Based Reinforcement Learning
- RadarMOSEVE: A Spatial-Temporal Transformer Network for Radar-Only Moving Object Segmentation and Ego-Velocity Estimation
- Efficient Reinforcement Learning for StarCraft by Abstract Forward Models and Transfer Learning
- Mega-Reward: Achieving Human-Level Play without Extrinsic Rewards
- Language-Mediated, Object-Centric Representation Learning
- Relevance-Guided Modeling of Object Dynamics for Reinforcement Learning
- Joint Perception and Control as Inference with an Object-based Implementation