Perspective Transformer Nets: Learning Single-View 3D Object Reconstruction without 3D Supervision
arXiv:1612.00814
Abstract
Understanding the 3D world is a fundamental problem in computer vision. However, learning a good representation of 3D objects is still an open problem due to the high dimensionality of the data and many factors of variation involved. In this work, we investigate the task of single-view 3D object reconstruction from a learning agent's perspective. We formulate the learning process as an interaction between 3D and 2D representations and propose an encoder-decoder network with a novel projection loss defined by the perspective transformation. More importantly, the projection loss enables the unsupervised learning using 2D observation without explicit 3D supervision. We demonstrate the ability of the model in generating 3D volume from a single 2D image with three sets of experiments: (1) learning from single-class objects; (2) learning from multi-class objects and (3) testing on novel object classes. Results show superior performance and better generalization ability for 3D object reconstruction when the projection loss is involved.
published at NIPS 2016
Cited by in corpus (105)
- Image-based 3D Object Reconstruction: State-of-the-Art and Trends in the Deep Learning Era
- Differentiable Surface Splatting for Point-based Geometry Processing
- DISN: Deep Implicit Surface Network for High-quality Single-view 3D Reconstruction
- Dense 3D Object Reconstruction from a Single Depth View
- Soft Rasterizer: A Differentiable Renderer for Image-based 3D Reasoning
- Differentiable Rendering: A Survey
- Soft Rasterizer: Differentiable Rendering for Unsupervised Single-View Mesh Reconstruction
- 3D Morphable Models as Spatial Transformer Networks
- Learning 3D Shape Completion under Weak Supervision
- Adversarial Attack and Defense on Point Sets
- RenderNet: A deep convolutional network for differentiable rendering from 3D shapes
- StructureNet: Hierarchical Graph Networks for 3D Shape Generation
- pixelNeRF: Neural Radiance Fields from One or Few Images
- NeRS: Neural Reflectance Surfaces for Sparse-view 3D Reconstruction in the Wild
- Deferred Neural Rendering: Image Synthesis using Neural Textures
- Weakly-Supervised Discovery of Geometry-Aware Representation for 3D Human Pose Estimation
- Lagrangian Neural Style Transfer for Fluids
- Pix2Vex: Image-to-Geometry Reconstruction using a Smooth Differentiable Renderer
- Implicit Geometric Regularization for Learning Shapes
- 3D-LMNet: Latent Embedding Matching for Accurate and Diverse 3D Point Cloud Reconstruction from a Single Image
- AIBench: An Industry Standard Internet Service AI Benchmark Suite
- DenseBody: Directly Regressing Dense 3D Human Pose and Shape From a Single Color Image
- Convolutional Generation of Textured 3D Meshes
- Data-Efficient Learning for Sim-to-Real Robotic Grasping using Deep Point Cloud Prediction Networks
- Learning Category-Specific Mesh Reconstruction from Image Collections
- HoliCity: A City-Scale Data Platform for Learning Holistic 3D Structures
- On Demand Solid Texture Synthesis Using Deep 3D Networks
- Compact Model Representation for 3D Reconstruction
- Monocular 3D Object Detection Leveraging Accurate Proposals and Shape Reconstruction
- DUP-Net: Denoiser and Upsampler Network for 3D Adversarial Point Clouds Defense
- Deep Mesh Reconstruction from Single RGB Images via Topology Modification Networks
- Implicit Mesh Reconstruction from Unannotated Image Collections
- SDF-SRN: Learning Signed Distance 3D Object Reconstruction from Static Images
- Geometric Capsule Autoencoders for 3D Point Clouds
- Learning Pose-invariant 3D Object Reconstruction from Single-view Images
- Neural Sparse Voxel Fields
- Higher-Order Function Networks for Learning Composable 3D Object Representations
- Deep-SLAM++: Object-level RGBD SLAM based on class-specific deep shape priors
- Unsupervised Discovery of 3D Physical Objects from Video
- SDFDiff: Differentiable Rendering of Signed Distance Fields for 3D Shape Optimization
- Multiview Detection with Feature Perspective Transformation
- Conditional Single-view Shape Generation for Multi-view Stereo Reconstruction
- 3DN: 3D Deformation Network
- Amodal 3D Reconstruction for Robotic Manipulation via Stability and Connectivity
- Localization and Mapping using Instance-specific Mesh Models
- Transformable Bottleneck Networks
- Semi-supervised Viewpoint Estimation with Geometry-aware Conditional Generation
- Variational Autoencoders for Deforming 3D Mesh Models
- PT2PC: Learning to Generate 3D Point Cloud Shapes from Part Tree Conditions
- Weakly Supervised Learning of Multi-Object 3D Scene Decompositions Using Deep Shape Priors
- Self-supervised Single-view 3D Reconstruction via Semantic Consistency
- Cooperative Holistic Scene Understanding: Unifying 3D Object, Layout, and Camera Pose Estimation
- Extreme Relative Pose Estimation for RGB-D Scans via Scene Completion
- ShapeAdv: Generating Shape-Aware Adversarial 3D Point Clouds
- Self-supervised Learning of 3D Objects from Natural Images
- Predicting Camera Viewpoint Improves Cross-dataset Generalization for 3D Human Pose Estimation
- Cerberus: A Multi-headed Derenderer
- MVPNet: Multi-View Point Regression Networks for 3D Object Reconstruction from A Single Image
- Generative Adversarial Frontal View to Bird View Synthesis
- Sketch2Model: View-Aware 3D Modeling from Single Free-Hand Sketches
- GSNet: Joint Vehicle Pose and Shape Reconstruction with Geometrical and Scene-aware Supervision
- Learning to Detect 3D Reflection Symmetry for Single-View Reconstruction
- TetraTSDF: 3D human reconstruction from a single image with a tetrahedral outer shell
- Learning with Algorithmic Supervision via Continuous Relaxations
- Learning Structural Graph Layouts and 3D Shapes for Long Span Bridges 3D Reconstruction
- Learning Manifold Patch-Based Representations of Man-Made Shapes
- Tracking by Animation: Unsupervised Learning of Multi-Object Attentive Trackers
- Layer-structured 3D Scene Inference via View Synthesis
- Learning Neural Light Transport
- AIBench Training: Balanced Industry-Standard AI Training Benchmarking
- SeqXY2SeqZ: Structure Learning for 3D Shapes by Sequentially Predicting 1D Occupancy Segments From 2D Coordinates
- GraphX-Convolution for Point Cloud Deformation in 2D-to-3D Conversion
- AIBench Scenario: Scenario-distilling AI Benchmarking
- SE-MD: A Single-encoder multiple-decoder deep network for point cloud generation from 2D images
- Pix2Surf: Learning Parametric 3D Surface Models of Objects from Images
- From Image Collections to Point Clouds with Self-supervised Shape and Pose Networks
- Dense 3D Point Cloud Reconstruction Using a Deep Pyramid Network
- RayNet: Learning Volumetric 3D Reconstruction with Ray Potentials
- StructEdit: Learning Structural Shape Variations
- PCLs: Geometry-aware Neural Reconstruction of 3D Pose with Perspective Crop Layers
- I-nteract 2.0: A Cyber-Physical System to Design 3D Models using Mixed Reality Technologies and Deep Learning for Additive Manufacturing
- Neural Articulated Radiance Field
- DR-KFS: A Differentiable Visual Similarity Metric for 3D Shape Reconstruction
- Neutral Face Game Character Auto-Creation via PokerFace-GAN
- Inferring 3D Shapes from Image Collections using Adversarial Networks
- AIBench: An Agile Domain-specific Benchmarking Methodology and an AI Benchmark Suite
- MeshMVS: Multi-View Stereo Guided Mesh Reconstruction
- Learning Equivariant Representations
- 3D Scattering Tomography by Deep Learning with Architecture Tailored to Cloud Fields
- Shelf-Supervised Mesh Prediction in the Wild
- Silhouette Guided Point Cloud Reconstruction beyond Occlusion
- Hierarchical View Predictor: Unsupervised 3D Global Feature Learning through Hierarchical Prediction among Unordered Views
- Neural Implicit 3D Shapes from Single Images with Spatial Patterns
- A Dataset-Dispersion Perspective on Reconstruction Versus Recognition in Single-View 3D Reconstruction Networks
- Learning Monocular 3D Vehicle Detection without 3D Bounding Box Labels
- Realtime Simulation of Thin-Shell Deformable Materials using CNN-Based Mesh Embedding
- Analytical Derivatives for Differentiable Renderer: 3D Pose Estimation by Silhouette Consistency
- 3D Ken Burns Effect from a Single Image
- Inverse Graphics: Unsupervised Learning of 3D Shapes from Single Images
- Learning to Reconstruct and Segment 3D Objects
- Fast and Explicit Neural View Synthesis
- Novel View Synthesis from a Single Image via Unsupervised learning
- Holodeck: Immersive 3D Displays Using Swarms of Flying Light Specks
- What Do Single-view 3D Reconstruction Networks Learn?
- GAMesh: Guided and Augmented Meshing for Deep Point Networks