Context-Aware Synthesis and Placement of Object Instances
arXiv:1812.02350
Abstract
Learning to insert an object instance into an image in a semantically coherent manner is a challenging and interesting problem. Solving it requires (a) determining a location to place an object in the scene and (b) determining its appearance at the location. Such an object insertion model can potentially facilitate numerous image editing and scene parsing applications. In this paper, we propose an end-to-end trainable neural network for the task of inserting an object instance mask of a specified class into the semantic label map of an image. Our network consists of two generative modules where one determines where the inserted object mask should be (i.e., location and scale) and the other determines what the object mask shape (and pose) should look like. The two modules are connected together via a spatial transformation network and jointly trained. We devise a learning procedure that leverage both supervised and unsupervised data and show our model can insert an object at diverse locations with various appearances. We conduct extensive experimental validations with comparisons to strong baselines to verify the effectiveness of the proposed network.
Cited by in corpus (24)
- SESAME: Semantic Editing of Scenes by Adding, Manipulating or Erasing Objects
- Data Augmentation for Object Detection via Progressive and Selective Instance-Switching
- Yes, we GAN: Applying Adversarial Techniques for Autonomous Driving
- Object Discovery with a Copy-Pasting GAN
- Making Images Real Again: A Comprehensive Survey on Deep Image Composition
- A Shape Transformation-based Dataset Augmentation Framework for Pedestrian Detection
- Mixing Real and Synthetic Data to Enhance Neural Network Training -- A Review of Current Approaches
- Towards Realistic 3D Embedding via View Alignment
- LayoutVAE: Stochastic Scene Layout Generation From a Label Set
- OPA: Object Placement Assessment Dataset
- AdvSPADE: Realistic Unrestricted Attacks for Semantic Segmentation
- Adversarial Image Composition with Auxiliary Illumination
- Long-term Human Motion Prediction with Scene Context
- BachGAN: High-Resolution Image Synthesis from Salient Object Layout
- Generating 3D People in Scenes without People
- GeoSim: Realistic Video Simulation via Geometry-Aware Composition for Self-Driving
- Putting Humans in a Scene: Learning Affordance in 3D Indoor Environments
- Controllable Image Synthesis via SegVAE
- GaussiGAN: Controllable Image Synthesis with 3D Gaussians from Unposed Silhouettes
- Road images augmentation with synthetic traffic signs using neural networks
- Inserting Videos into Videos
- An Unpaired Shape Transforming Method for Image Translation and Cross-Domain Retrieval
- Hidden Footprints: Learning Contextual Walkability from 3D Human Trails
- Multimodal Image Outpainting With Regularized Normalized Diversification