Render for CNN: Viewpoint Estimation in Images Using CNNs Trained with Rendered 3D Model Views
arXiv:1505.05641
Abstract
Object viewpoint estimation from 2D images is an essential task in computer vision. However, two issues hinder its progress: scarcity of training data with viewpoint annotations, and a lack of powerful features. Inspired by the growing availability of 3D models, we propose a framework to address both issues by combining render-based image synthesis and CNNs. We believe that 3D models have the potential in generating a large number of images of high variation, which can be well exploited by deep CNN with a high learning capacity. Towards this goal, we propose a scalable and overfit-resistant image synthesis pipeline, together with a novel CNN specifically tailored for the viewpoint estimation task. Experimentally, we show that the viewpoint estimation from our pipeline can significantly outperform state-of-the-art methods on PASCAL 3D+ benchmark.
References in corpus (4)
Cited by in corpus (25)
- Domain Adaptation for Visual Applications: A Comprehensive Survey
- Domain Randomization for Transferring Deep Neural Networks from Simulation to the Real World
- MoCap-guided Data Augmentation for 3D Pose Estimation in the Wild
- Object Detection Using Deep CNNs Trained on Synthetic Images
- Rethinking Reprojection: Closing the Loop for Pose-aware ShapeReconstruction from a Single Image
- Transformation-Grounded Image Generation Network for Novel 3D View Synthesis
- SurfNet: Generating 3D shape surfaces using deep residual networks
- 3D Shape Induction from 2D Views of Multiple Objects
- Deep Cuboid Detection: Beyond 2D Bounding Boxes
- Crafting a multi-task CNN for viewpoint estimation
- FacePoseNet: Making a Case for Landmark-Free Face Alignment
- Synthetic to Real Adaptation with Generative Correlation Alignment Networks
- Visual Compiler: Synthesizing a Scene-Specific Pedestrian Detector and Pose Estimator
- 3D Reconstruction of Simple Objects from A Single View Silhouette Image
- Shape Generation using Spatially Partitioned Point Clouds
- Pano2CAD: Room Layout From A Single Panorama Image
- A spatiotemporal model with visual attention for video classification
- Improved Deep Learning of Object Category using Pose Information
- Object-Centric Photometric Bundle Adjustment with Deep Shape Prior
- Shape from Shading through Shape Evolution
- Learning to Recognize Objects by Retaining other Factors of Variation
- A Dataset for Developing and Benchmarking Active Vision
- Synthesizing 3D Shapes from Silhouette Image Collections using Multi-projection Generative Adversarial Networks
- Fast 3D Pose Refinement with RGB Images
- From Virtual to Real World Visual Perception using Domain Adaptation -- The DPM as Example