The Replica Dataset: A Digital Replica of Indoor Spaces
arXiv:1906.05797
Abstract
We introduce Replica, a dataset of 18 highly photo-realistic 3D indoor scene reconstructions at room and building scale. Each scene consists of a dense mesh, high-resolution high-dynamic-range (HDR) textures, per-primitive semantic class and instance information, and planar mirror and glass reflectors. The goal of Replica is to enable machine learning (ML) research that relies on visually, geometrically, and semantically realistic generative models of the world - for instance, egocentric computer vision, semantic segmentation in 2D and 3D, geometric inference, and the development of embodied agents (virtual robots) performing navigation, instruction following, and question answering. Due to the high level of realism of the renderings from Replica, there is hope that ML systems trained on Replica may transfer directly to real world image and video data. Together with the data, we are releasing a minimal C++ SDK as a starting point for working with the Replica dataset. In addition, Replica is `Habitat-compatible', i.e. can be natively used with AI Habitat for training and testing embodied agents.
Cited by in corpus (50)
- Orbit: A Unified Simulation Framework for Interactive Robot Learning Environments
- Vox-Fusion: Dense Tracking and Mapping with Voxel-based Neural Implicit Representation
- ROSEFusion: Random Optimization for Online Dense Reconstruction under Fast Camera Motion
- Vision-and-Language Navigation: A Survey of Tasks, Methods, and Future Directions
- PIN-SLAM: LiDAR SLAM Using a Point-Based Implicit Neural Representation for Achieving Global Map Consistency
- Gaussian Splatting: 3D Reconstruction and Novel View Synthesis, a Review
- SceneDreamer: Unbounded 3D Scene Generation from 2D Image Collections
- H2-Mapping: Real-time Dense Mapping Using Hierarchical Hybrid Representation
- Habitat-Matterport 3D Dataset (HM3D): 1000 Large-scale 3D Environments for Embodied AI
- A Survey of Embodied AI: From Simulators to Research Tasks
- Habitat 2.0: Training Home Assistants to Rearrange their Habitat
- Evaluating Modern Approaches in 3D Scene Reconstruction: NeRF vs Gaussian-Based Methods
- RO-MAP: Real-Time Multi-Object Mapping with Neural Radiance Fields
- Towards Open World NeRF-Based SLAM
- ClearGrasp: 3D Shape Estimation of Transparent Objects for Manipulation
- Benchmarking Neural Radiance Fields for Autonomous Robots: An Overview
- ILabel: Interactive Neural Scene Labelling
- Move to See Better: Self-Improving Embodied Object Detection
- O2O-Afford: Annotation-Free Large-Scale Object-Object Affordance Learning
- In-Place Scene Labelling and Understanding with Implicit Scene Representation
- Learning Semantic-Agnostic and Spatial-Aware Representation for Generalizable Visual-Audio Navigation
- NVS-MonoDepth: Improving Monocular Depth Prediction with Novel View Synthesis
- Panoptic Vision-Language Feature Fields
- Simultaneous Localization and Mapping Related Datasets: A Comprehensive Survey
- VisualEchoes: Spatial Image Representation Learning through Echolocation
- Data Augmentation for Object Detection via Differentiable Neural Rendering
- HAMMER: Heterogeneous, Multi-Robot Semantic Gaussian Splatting
- Unsupervised Domain Adaptation for Visual Navigation
- Learning Audio-Visual Dereverberation
- Environment Predictive Coding for Embodied Agents
- Recognizing Scenes from Novel Viewpoints
- NeSF: Neural Semantic Fields for Generalizable Semantic Segmentation of 3D Scenes
- Bonn Activity Maps: Dataset Description
- AiSDF: Structure-aware Neural Signed Distance Fields in Indoor Scenes
- Semantic Is Enough: Only Semantic Information For NeRF Reconstruction
- Scan2Part: Fine-grained and Hierarchical Part-level Understanding of Real-World 3D Scans
- Robustness via Cross-Domain Ensembles
- OV-MAP: Open-Vocabulary Zero-Shot 3D Instance Segmentation Map for Robots
- ConfidentSplat: Confidence-Weighted Depth Fusion for Accurate 3D Gaussian Splatting SLAM
- The Surprising Effectiveness of Visual Odometry Techniques for Embodied PointGoal Navigation
- 3D Object Recognition By Corresponding and Quantizing Neural 3D Scene Representations
- Egocentric Activity Recognition and Localization on a 3D Map
- Semantic Audio-Visual Navigation
- Learning to compose 6-DoF omnidirectional videos using multi-sphere images
- Localising In Complex Scenes Using Balanced Adversarial Adaptation
- Resolution Where It Counts: Hash-based GPU-Accelerated 3D Reconstruction via Variance-Adaptive Voxel Grids
- Building Intelligent Autonomous Navigation Agents
- MINERVAS: Massive INterior EnviRonments VirtuAl Synthesis
- A Survey on Deep Learning Architectures for Point Cloud Classification and Segmentation
- Video Autoencoder: self-supervised disentanglement of static 3D structure and motion