Building Generalizable Agents with a Realistic and Rich 3D Environment
arXiv:1801.02209
Abstract
Teaching an agent to navigate in an unseen 3D environment is a challenging task, even in the event of simulated environments. To generalize to unseen environments, an agent needs to be robust to low-level variations (e.g. color, texture, object changes), and also high-level variations (e.g. layout changes of the environment). To improve overall generalization, all types of variations in the environment have to be taken under consideration via different level of data augmentation steps. To this end, we propose House3D, a rich, extensible and efficient environment that contains 45,622 human-designed 3D scenes of visually realistic houses, ranging from single-room studios to multi-storied houses, equipped with a diverse set of fully labeled 3D objects, textures and scene layouts, based on the SUNCG dataset (Song et.al.). The diversity in House3D opens the door towards scene-level augmentation, while the label-rich nature of House3D enables us to inject pixel- & task-level augmentations such as domain randomization (Toubin et. al.) and multi-task training. Using a subset of houses in House3D, we show that reinforcement learning agents trained with an enhancement of different levels of augmentations perform much better in unseen environments than our baselines with raw RGB input by over 8% in terms of navigation success rate. House3D is publicly available at http://github.com/facebookresearch/House3D.
updated with improved content and more experinemnts
Cited by in corpus (84)
- On Evaluation of Embodied Navigation Agents
- AI2-THOR: An Interactive 3D Environment for Visual AI
- Cross-view Semantic Segmentation for Sensing Surroundings
- Dark, Beyond Deep: A Paradigm Shift to Cognitive AI with Humanlike Common Sense
- ThreeDWorld: A Platform for Interactive Multi-Modal Physical Simulation
- Vision-and-Language Navigation: A Survey of Tasks, Methods, and Future Directions
- ViZDoom Competitions: Playing Doom from Pixels
- Spatial Action Maps for Mobile Manipulation
- CHALET: Cornell House Agent Learning Environment
- Domain Stylization: A Strong, Simple Baseline for Synthetic to Real Image Domain Adaptation
- Learning to Navigate in Cities Without a Map
- Deep Learning Based 3D Segmentation: A Survey
- A Comprehensive Review of Modern Object Segmentation Approaches
- Watch-And-Help: A Challenge for Social Perception and Human-AI Collaboration
- Benchmarking Classic and Learned Navigation in Complex 3D Environments
- Emergence of Exploratory Look-Around Behaviors through Active Observation Completion
- Learning Exploration Policies for Navigation
- VRKitchen: an Interactive 3D Virtual Environment for Task-oriented Learning
- Blindfold Baselines for Embodied QA
- Human Instruction-Following with Deep Reinforcement Learning via Transfer-Learning from Text
- ManiSkill: Generalizable Manipulation Skill Benchmark with Large-Scale Demonstrations
- Habitat 2.0: Training Home Assistants to Rearrange their Habitat
- Guided Feature Transformation (GFT): A Neural Language Grounding Module for Embodied Agents
- Improving Target-driven Visual Navigation with Attention on 3D Spatial Relationships
- Embodied Visual Recognition
- A Behavioral Approach to Visual Navigation with Graph Localization Networks
- Adversarial Reinforced Instruction Attacker for Robust Vision-Language Navigation
- Particle Filter Networks with Application to Visual Localization
- Collaborative Visual Navigation
- VirtualHome: Simulating Household Activities via Programs
- A Mask-RCNN Baseline for Probabilistic Object Detection
- From Seeing to Moving: A Survey on Learning for Visual Indoor Navigation (VIN)
- The Regretful Agent: Heuristic-Aided Navigation through Progress Estimation
- Embodied Question Answering in Photorealistic Environments with Point Cloud Perception
- REVERIE: Remote Embodied Visual Referring Expression in Real Indoor Environments
- 3DB: A Framework for Debugging Computer Vision Models
- Evolutionary Population Curriculum for Scaling Multi-Agent Reinforcement Learning
- What Should I Do Now? Marrying Reinforcement Learning and Symbolic Planning
- Discovering Diverse Multi-Agent Strategic Behavior via Reward Randomization
- OpenRooms: An End-to-End Open Framework for Photorealistic Indoor Scene Datasets
- Mis-spoke or mis-lead: Achieving Robustness in Multi-Agent Communicative Reinforcement Learning
- Learning and Planning with a Semantic Model
- VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering
- Why Build an Assistant in Minecraft?
- The ThreeDWorld Transport Challenge: A Visually Guided Task-and-Motion Planning Benchmark for Physically Realistic Embodied AI
- Answering Visual What-If Questions: From Actions to Predicted Scene Descriptions
- Learning to Move with Affordance Maps
- Arena: A General Evaluation Platform and Building Toolkit for Multi-Agent Intelligence
- Sim-Real Joint Reinforcement Transfer for 3D Indoor Navigation
- Bayesian Relational Memory for Semantic Visual Navigation
- Learning to Generate Synthetic Data via Compositing
- Unsupervised Reinforcement Learning of Transferable Meta-Skills for Embodied Navigation
- End-to-end Active Object Tracking and Its Real-world Deployment via Reinforcement Learning
- Unsupervised Domain Adaptation for Visual Navigation
- Scalable Modular Synthetic Data Generation for Advancing Aerial Autonomy
- Mastering emergent language: learning to guide in simulated navigation
- Influence-Based Multi-Agent Exploration
- SurfelGAN: Synthesizing Realistic Sensor Data for Autonomous Driving
- Visual Hindsight Self-Imitation Learning for Interactive Navigation
- Scene-Intuitive Agent for Remote Embodied Visual Grounding
- Learning to Map for Active Semantic Goal Navigation
- Predicting Performance of SLAM Algorithms
- Reinforcement Learning-based Visual Navigation with Information-Theoretic Regularization
- Zero-Shot Compositional Policy Learning via Language Grounding
- Learning to simulate complex scenes
- Explore the Potential Performance of Vision-and-Language Navigation Model: a Snapshot Ensemble Method
- Walking with MIND: Mental Imagery eNhanceD Embodied QA
- Continual Reinforcement Learning with Diversity Exploration and Adversarial Self-Correction
- Generative Language-Grounded Policy in Vision-and-Language Navigation with Bayes' Rule
- NeoNav: Improving the Generalization of Visual Navigation via Generating Next Expected Observations
- PHYRE: A New Benchmark for Physical Reasoning
- Efficient Robotic Object Search via HIEM: Hierarchical Policy Learning with Intrinsic-Extrinsic Modeling
- GAPLE: Generalizable Approaching Policy LEarning for Robotic Object Searching in Indoor Environment
- The Surprising Effectiveness of Visual Odometry Techniques for Embodied PointGoal Navigation
- Are you doing what I say? On modalities alignment in ALFRED
- Hierarchical Cross-Modal Agent for Robotics Vision-and-Language Navigation
- Exploiting Language Instructions for Interpretable and Compositional Reinforcement Learning
- Hierarchical and Partially Observable Goal-driven Policy Learning with Goals Relational Graph
- A Self-Supervised Auxiliary Loss for Deep RL in Partially Observable Settings
- MINERVAS: Massive INterior EnviRonments VirtuAl Synthesis
- Multiplicative Gaussian Particle Filter
- Reinforced Natural Language Interfaces via Entropy Decomposition
- Language coverage and generalization in RNN-based continuous sentence embeddings for interacting agents
- Learning to Navigate Unseen Environments: Back Translation with Environmental Dropout