AI2-THOR: An Interactive 3D Environment for Visual AI
arXiv:1712.05474
Abstract
We introduce The House Of inteRactions (THOR), a framework for visual AI research, available at http://ai2thor.allenai.org. AI2-THOR consists of near photo-realistic 3D indoor scenes, where AI agents can navigate in the scenes and interact with objects to perform tasks. AI2-THOR enables research in many different domains including but not limited to deep reinforcement learning, imitation learning, learning by interaction, planning, visual question answering, unsupervised representation learning, object detection and segmentation, and learning models of cognition. The goal of AI2-THOR is to facilitate building visually intelligent models and push the research forward in this domain.
References in corpus (7)
- Matterport3D: Learning from RGB-D Data in Indoor Environments
- Building Generalizable Agents with a Realistic and Rich 3D Environment
- MINOS: Multimodal Indoor Simulator for Navigation in Complex Environments
- SceneNet RGB-D: 5M Photorealistic Images of Synthetic Indoor Trajectories with Ground Truth
- TorchCraft: a Library for Machine Learning Research on Real-Time Strategy Games
- HoME: a Household Multimodal Environment
- DeepMind Lab
Cited by in corpus (140)
- On the Opportunities and Risks of Foundation Models
- Unity: A General Platform for Intelligent Agents
- On Evaluation of Embodied Navigation Agents
- Meta-World: A Benchmark and Evaluation for Multi-Task and Meta Reinforcement Learning
- robosuite: A Modular Simulation Framework and Benchmark for Robot Learning
- Dark, Beyond Deep: A Paradigm Shift to Cognitive AI with Humanlike Common Sense
- ThreeDWorld: A Platform for Interactive Multi-Modal Physical Simulation
- Vision-and-Dialog Navigation
- Rearrangement: A Challenge for Embodied AI
- ViZDoom Competitions: Playing Doom from Pixels
- Spatial Action Maps for Mobile Manipulation
- Unity Perception: Generate Synthetic Data for Computer Vision
- iGibson 2.0: Object-Centric Simulation for Robot Learning of Everyday Household Tasks
- Imitating Latent Policies from Observation
- Learning to Navigate in Cities Without a Map
- The StreetLearn Environment and Dataset
- AllenAct: A Framework for Embodied AI Research
- EvalAI: Towards Better Evaluation Systems for AI Agents
- iGibson 1.0: a Simulation Environment for Interactive Tasks in Large Realistic Scenes
- Watch-And-Help: A Challenge for Social Perception and Human-AI Collaboration
- PlaTe: Visually-Grounded Planning with Transformers in Procedural Tasks
- Reinforced Cross-Modal Matching and Self-Supervised Imitation Learning for Vision-Language Navigation
- BEHAVIOR: Benchmark for Everyday Household Activities in Virtual, Interactive, and Ecological Environments
- Habitat-Matterport 3D Dataset (HM3D): 1000 Large-scale 3D Environments for Embodied AI
- VTNet: Visual Transformer Network for Object Goal Navigation
- VRKitchen: an Interactive 3D Virtual Environment for Task-oriented Learning
- Simulating Content Consistent Vehicle Datasets with Attribute Descent
- Learning hierarchical relationships for object-goal navigation
- Blindfold Baselines for Embodied QA
- ALFWorld: Aligning Text and Embodied Environments for Interactive Learning
- Natural Environment Benchmarks for Reinforcement Learning
- ManiSkill: Generalizable Manipulation Skill Benchmark with Large-Scale Demonstrations
- A Survey of Embodied AI: From Simulators to Research Tasks
- Habitat 2.0: Training Home Assistants to Rearrange their Habitat
- The AdobeIndoorNav Dataset: Towards Deep Reinforcement Learning based Real-world Indoor Robot Visual Navigation
- LanguageRefer: Spatial-Language Model for 3D Visual Grounding
- Embodied Visual Recognition
- Improving Target-driven Visual Navigation with Attention on 3D Spatial Relationships
- Domain Randomization and Pyramid Consistency: Simulation-to-Real Generalization without Accessing Target Domain Data
- Learning Object Relation Graph and Tentative Policy for Visual Navigation
- Collaborative Visual Navigation
- CraftAssist: A Framework for Dialogue-enabled Interactive Agents
- Towards Learning a Generic Agent for Vision-and-Language Navigation via Pre-training
- Baby Intuitions Benchmark (BIB): Discerning the goals, preferences, and actions of others
- CATER: A diagnostic dataset for Compositional Actions and TEmporal Reasoning
- A Mask-RCNN Baseline for Probabilistic Object Detection
- Vision-Language Navigation with Self-Supervised Auxiliary Reasoning Tasks
- VirtualHome: Simulating Household Activities via Programs
- Neural Task Graphs: Generalizing to Unseen Tasks from a Single Video Demonstration
- From Seeing to Moving: A Survey on Learning for Visual Indoor Navigation (VIN)
- FILM: Following Instructions in Language with Modular Methods
- Embodied Question Answering in Photorealistic Environments with Point Cloud Perception
- The Regretful Agent: Heuristic-Aided Navigation through Progress Estimation
- 3DB: A Framework for Debugging Computer Vision Models
- REVERIE: Remote Embodied Visual Referring Expression in Real Indoor Environments
- What Should I Do Now? Marrying Reinforcement Learning and Symbolic Planning
- OpenRooms: An End-to-End Open Framework for Photorealistic Indoor Scene Datasets
- Deep Learning for Embodied Vision Navigation: A Survey
- Learning Generalizable Visual Representations via Interactive Gameplay
- Cross-Domain Facial Expression Recognition: A Unified Evaluation Benchmark and Adversarial Graph Learning
- Learning Affordance Landscapes for Interaction Exploration in 3D Environments
- Synthesized Policies for Transfer and Adaptation across Tasks and Environments
- Multi-View Learning for Vision-and-Language Navigation
- Evaluating the Impact of Semantic Segmentation and Pose Estimation on Dense Semantic SLAM
- Why Build an Assistant in Minecraft?
- VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering
- Look, Listen, and Act: Towards Audio-Visual Embodied Navigation
- The ThreeDWorld Transport Challenge: A Visually Guided Task-and-Motion Planning Benchmark for Physically Realistic Embodied AI
- CORA: Benchmarks, Baselines, and Metrics as a Platform for Continual Reinforcement Learning Agents
- BenchBot: Evaluating Robotics Research in Photorealistic 3D Simulation and on Real Robots
- NViSII: A Scriptable Tool for Photorealistic Image Generation
- Multi-Target Embodied Question Answering
- Arena: A General Evaluation Platform and Building Toolkit for Multi-Agent Intelligence
- Fever Basketball: A Complex, Flexible, and Asynchronized Sports Game Environment for Multi-agent Reinforcement Learning
- Are We There Yet? Learning to Localize in Embodied Instruction Following
- Sim-Real Joint Reinforcement Transfer for 3D Indoor Navigation
- No RL, No Simulation: Learning to Navigate without Navigating
- SoundSpaces: Audio-Visual Navigation in 3D Environments
- Bayesian Relational Memory for Semantic Visual Navigation
- Unsupervised Domain Adaptation for Visual Navigation
- Hierarchical Task Learning from Language Instructions with Unified Transformers and Self-Monitoring
- An A* Curriculum Approach to Reinforcement Learning for RGBD Indoor Robot Navigation
- SAILenv: Learning in Virtual Visual Environments Made Simple
- Unsupervised Reinforcement Learning of Transferable Meta-Skills for Embodied Navigation
- Embodied BERT: A Transformer Model for Embodied, Language-guided Visual Task Completion
- Navigation Agents for the Visually Impaired: A Sidewalk Simulator and Experiments
- Discovery of Options via Meta-Learned Subgoals
- Environment Predictive Coding for Embodied Agents
- SASRA: Semantically-aware Spatio-temporal Reasoning Agent for Vision-and-Language Navigation in Continuous Environments
- Learning About Objects by Learning to Interact with Them
- LUMINOUS: Indoor Scene Generation for Embodied AI Challenges
- The Robotic Vision Scene Understanding Challenge
- Target Driven Visual Navigation with Hybrid Asynchronous Universal Successor Representations
- GeoSim: Realistic Video Simulation via Geometry-Aware Composition for Self-Driving
- Learning to Map for Active Semantic Goal Navigation
- Large Batch Simulation for Deep Reinforcement Learning
- ArraMon: A Joint Navigation-Assembly Instruction Interpretation Task in Dynamic Environments
- Bridging the Imitation Gap by Adaptive Insubordination
- Adaptive Procedural Task Generation for Hard-Exploration Problems
- Zero-Shot Compositional Policy Learning via Language Grounding
- Vision-Dialog Navigation by Exploring Cross-modal Memory
- Shaping embodied agent behavior with activity-context priors from egocentric video
- Meta Reinforcement Learning with Autonomous Inference of Subtask Dependencies
- Explore the Potential Performance of Vision-and-Language Navigation Model: a Snapshot Ensemble Method
- Learning to Generate Synthetic 3D Training Data through Hybrid Gradient
- Learning to simulate complex scenes
- Cross-View Policy Learning for Street Navigation
- Dynamic Value Estimation for Single-Task Multi-Scene Reinforcement Learning
- Optimal Assistance for Object-Rearrangement Tasks in Augmented Reality
- Landmark Policy Optimization for Object Navigation Task
- Learning Adaptive Language Interfaces through Decomposition
- MaAST: Map Attention with Semantic Transformersfor Efficient Visual Navigation
- Bridging Scene Understanding and Task Execution with Flexible Simulation Environments
- Sparse Attention Guided Dynamic Value Estimation for Single-Task Multi-Scene Reinforcement Learning
- NavigationNet: A Large-scale Interactive Indoor Navigation Dataset
- SIGVerse: A cloud-based VR platform for research on social and embodied human-robot interaction
- SeanNet: Semantic Understanding Network for Localization Under Object Dynamics
- Walking with MIND: Mental Imagery eNhanceD Embodied QA
- Dialogue Object Search
- VRGym: A Virtual Testbed for Physical and Interactive AI
- Hierarchical Cross-Modal Agent for Robotics Vision-and-Language Navigation
- Multi-Agent Embodied Visual Semantic Navigation with Scene Prior Knowledge
- Multimodal Aggregation Approach for Memory Vision-Voice Indoor Navigation with Meta-Learning
- Discovering Generalizable Skills via Automated Generation of Diverse Tasks
- Learning Autonomous Exploration and Mapping with Semantic Vision
- Leveraging Semantics for Incremental Learning in Multi-Relational Embeddings
- MINERVAS: Massive INterior EnviRonments VirtuAl Synthesis
- Reinforcement Learning for Sparse-Reward Object-Interaction Tasks in a First-person Simulated 3D Environment
- Reconstructing Interactive 3D Scenes by Panoptic Mapping and CAD Model Alignments
- Enhanced Scene Specificity with Sparse Dynamic Value Estimation
- VSGM -- Enhance robot task understanding ability through visual semantic graph
- Language coverage and generalization in RNN-based continuous sentence embeddings for interacting agents
- Messing Up 3D Virtual Environments: Transferable Adversarial 3D Objects
- Active Object Perceiver: Recognition-guided Policy Learning for Object Searching on Mobile Robots
- Contextual Scene Augmentation and Synthesis via GSACNet
- Hierarchical and Partially Observable Goal-driven Policy Learning with Goals Relational Graph
- Spatiotemporal Attacks for Embodied Agents
- Pushing it out of the Way: Interactive Visual Navigation
- Synthetic Data Are as Good as the Real for Association Knowledge Learning in Multi-object Tracking
- Robot in a China Shop: Using Reinforcement Learning for Location-Specific Navigation Behaviour