Reinforcement Learning-based Visual Navigation with Information-Theoretic Regularization
arXiv:1912.04078
Abstract
To enhance the cross-target and cross-scene generalization of target-driven visual navigation based on deep reinforcement learning (RL), we introduce an information-theoretic regularization term into the RL objective. The regularization maximizes the mutual information between navigation actions and visual observation transforms of an agent, thus promoting more informed navigation decisions. This way, the agent models the action-observation dynamics by learning a variational generative model. Based on the model, the agent generates (imagines) the next observation from its current observation and navigation target. This way, the agent learns to understand the causality between navigation actions and the changes in its observations, which allows the agent to predict the next action for navigation by comparing the current and the imagined next observations. Cross-target and cross-scene evaluations on the AI2-THOR framework show that our method attains at least a improvement of average success rate over some state-of-the-art models. We further evaluate our model in two real-world settings: navigation in unseen indoor scenes from a discrete Active Vision Dataset (AVD) and continuous real-world environments with a TurtleBot.We demonstrate that our navigation model is able to successfully achieve navigation tasks in these scenarios. Videos and models can be found in the supplementary material.
corresponding author: Kai Xu ([email protected]) and Jun Wang ([email protected]), accepted by IEEE Robotics and Automation Letters
References in corpus (13)
- Spectral Normalization for Generative Adversarial Networks
- Self-Attention Generative Adversarial Networks
- Recurrent World Models Facilitate Policy Evolution
- Building Generalizable Agents with a Realistic and Rich 3D Environment
- Object Goal Navigation using Goal-Oriented Semantic Exploration
- Reinforcement and Imitation Learning for Diverse Visuomotor Skills
- Reinforcement Learning from Imperfect Demonstrations
- Deeply AggreVaTeD: Differentiable Imitation Learning for Sequential Prediction
- Learning model-based planning from scratch
- Semi-parametric Topological Memory for Navigation
- Visual Semantic Navigation using Scene Priors
- A Regularized Approach to Sparse Optimal Policy in Reinforcement Learning
- Policy Optimization Reinforcement Learning with Entropy Regularization