Incremental Scene Synthesis
arXiv:1811.12297
Abstract
We present a method to incrementally generate complete 2D or 3D scenes with the following properties: (a) it is globally consistent at each step according to a learned scene prior, (b) real observations of a scene can be incorporated while observing global consistency, (c) unobserved regions can be hallucinated locally in consistence with previous observations, hallucinations and global priors, and (d) hallucinations are statistical in nature, i.e., different scenes can be generated from the same observations. To achieve this, we model the virtual scene, where an active agent at each step can either perceive an observed part of the scene or generate a local hallucination. The latter can be interpreted as the agent's expectation at this step through the scene and can be applied to autonomous navigation. In the limit of observing real data at each point, our method converges to solving the SLAM problem. It can otherwise sample entirely imagined scenes from prior distributions. Besides autonomous agents, applications include problems where large data is required for building robust real-world applications, but few samples are available. We demonstrate efficacy on various 2D as well as 3D data.
References in corpus (17)
- Adam: A Method for Stochastic Optimization
- Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks
- Spectral Normalization for Generative Adversarial Networks
- Self-Attention Generative Adversarial Networks
- Deep Convolutional Inverse Graphics Network
- Multi-Task Learning Using Uncertainty to Weigh Losses for Scene Geometry and Semantics
- Long Short-Term Memory-Networks for Machine Reading
- Neural SLAM: Learning to Explore with External Memory
- ViZDoom Competitions: Playing Doom from Pixels
- DeepStereo: Learning to Predict New Views from the World's Imagery
- HoME: a Household Multimodal Environment
- CNN-SLAM: Real-time dense monocular SLAM with learned depth prediction
- PAD-Net: Multi-Tasks Guided Prediction-and-Distillation Network for Simultaneous Depth Estimation and Scene Parsing
- Semantic Scene Completion from a Single Depth Image
- Transformation-Grounded Image Generation Network for Novel 3D View Synthesis
- Active Neural Localization
- Learning models for visual 3D localization with implicit mapping