Playing for Data: Ground Truth from Computer Games
arXiv:1608.02192
Abstract
Recent progress in computer vision has been driven by high-capacity models trained on large datasets. Unfortunately, creating large datasets with pixel-level labels has been extremely costly due to the amount of human effort required. In this paper, we present an approach to rapidly creating pixel-accurate semantic label maps for images extracted from modern computer games. Although the source code and the internal operation of commercial games are inaccessible, we show that associations between image patches can be reconstructed from the communication between the game and the graphics hardware. This enables rapid propagation of semantic labels within and across images synthesized by the game, with no access to the source code or the content. We validate the presented approach by producing dense pixel-level semantic annotations for 25 thousand images synthesized by a photorealistic open-world computer game. Experiments on semantic segmentation datasets show that using the acquired data to supplement real-world images significantly increases accuracy and that the acquired data enables reducing the amount of hand-labeled real-world data: models trained with game data and just 1/3 of the CamVid training set outperform models trained on the complete CamVid training set.
Accepted to the 14th European Conference on Computer Vision (ECCV 2016)
References in corpus (1)
Cited by in corpus (19)
- FCNs in the Wild: Pixel-level Adversarial and Constraint-based Adaptation
- CyCADA: Cycle-Consistent Adversarial Domain Adaptation
- Domain Adaptation for Visual Applications: A Comprehensive Survey
- CARLA: An Open Urban Driving Simulator
- Domain Randomization for Transferring Deep Neural Networks from Simulation to the Real World
- SceneNet RGB-D: 5M Photorealistic Images of Synthetic Indoor Trajectories with Ground Truth
- SqueezeSeg: Convolutional Neural Nets with Recurrent CRF for Real-Time Road-Object Segmentation from 3D LiDAR Point Cloud
- Unsupervised Histopathology Image Synthesis
- Image to Image Translation for Domain Adaptation
- TorontoCity: Seeing the World with a Million Eyes
- Synthetic to Real Adaptation with Generative Correlation Alignment Networks
- LabelFusion: A Pipeline for Generating Ground Truth Labels for Real RGBD Data of Cluttered Scenes
- The ParallelEye Dataset: Constructing Large-Scale Artificial Scenes for Traffic Vision Research
- SAD-GAN: Synthetic Autonomous Driving using Generative Adversarial Networks
- Cross-Domain Self-supervised Multi-task Feature Learning using Synthetic Imagery
- Training and Testing Object Detectors with Virtual Images
- Shape from Shading through Shape Evolution
- 3D Trajectory Reconstruction of Dynamic Objects Using Planarity Constraints
- Learning to Map Vehicles into Bird's Eye View