Multi-view Convolutional Neural Networks for 3D Shape Recognition
arXiv:1505.00880
Abstract
A longstanding question in computer vision concerns the representation of 3D shapes for recognition: should 3D shapes be represented with descriptors operating on their native 3D formats, such as voxel grid or polygon mesh, or can they be effectively represented with view-based descriptors? We address this question in the context of learning to recognize 3D shapes from a collection of their rendered views on 2D images. We first present a standard CNN architecture trained to recognize the shapes' rendered views independently of each other, and show that a 3D shape can be recognized even from a single view at an accuracy far higher than using state-of-the-art 3D shape descriptors. Recognition rates further increase when multiple views of the shapes are provided. In addition, we present a novel CNN architecture that combines information from multiple views of a 3D shape into a single and compact shape descriptor offering even better recognition performance. The same architecture can be applied to accurately recognize human hand-drawn sketches of shapes. We conclude that a collection of 2D views can be highly informative for 3D shape recognition and is amenable to emerging CNN architectures and their derivatives.
v1: Initial version. v2: An updated ModelNet40 training/test split is used; results with low-rank Mahalanobis metric learning are added. v3 (ICCV 2015): A second camera setup without the upright orientation assumption is added; some accuracy and mAP numbers are changed slightly because a small issue in mesh rendering related to specularities is fixed
References in corpus (3)
Cited by in corpus (71)
- Geometric deep learning: going beyond Euclidean data
- PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation
- PointNet++: Deep Hierarchical Feature Learning on Point Sets in a Metric Space
- PointSIFT: A SIFT-like Network Module for 3D Point Cloud Semantic Segmentation
- Dynamic Graph CNN for Learning on Point Clouds
- Directionally Constrained Fully Convolutional Neural Network For Airborne Lidar Point Cloud Classification
- Unsupervised Learning of 3D Structure from Images
- Do We Really Need to Collect Millions of Faces for Effective Face Recognition?
- Density-Aware Convolutional Networks with Context Encoding for Airborne LiDAR Point Cloud Classification
- Local Spectral Graph Convolution for Point Set Feature Learning
- ShellNet: Efficient Point Cloud Convolutional Neural Networks using Concentric Shells Statistics
- Dynamic Edge-Conditioned Filters in Convolutional Neural Networks on Graphs
- 3D Point Cloud Classification and Segmentation using 3D Modified Fisher Vector Representation for Convolutional Neural Networks
- Tangent Convolutions for Dense Prediction in 3D
- Point2Sequence: Learning the Shape Representation of 3D Point Clouds with an Attention-based Sequence to Sequence Network
- OctNet: Learning Deep 3D Representations at High Resolutions
- Spatio-temporal Stacked LSTM for Temperature Prediction in Weather Forecasting
- PVNet: A Joint Convolutional Network of Point Cloud and Multi-View for 3D Shape Recognition
- Noise-resistant Deep Learning for Object Classification in 3D Point Clouds Using a Point Pair Descriptor
- PyramNet: Point Cloud Pyramid Attention Network and Graph Embedding Module for Classification and Segmentation
- LPD-Net: 3D Point Cloud Learning for Large-Scale Place Recognition and Environment Analysis
- Interactive 3D Modeling with a Generative Adversarial Network
- Weakly Supervised Semantic Segmentation in 3D Graph-Structured Point Clouds of Wild Scenes
- PFCNN: Convolutional Neural Networks on 3D Surfaces Using Parallel Frames
- Morphological Error Detection in 3D Segmentations
- Multi-View Deep Learning for Consistent Semantic Mapping with RGB-D Cameras
- Shallow2Deep: Indoor Scene Modeling by Single Image Understanding
- DAPnet: A Double Self-attention Convolutional Network for Point Cloud Semantic Labeling
- Modeling Local Geometric Structure of 3D Point Clouds using Geo-CNN
- V2VNet: Vehicle-to-Vehicle Communication for Joint Perception and Prediction
- Spatial Aggregation of Holistically-Nested Convolutional Neural Networks for Automated Pancreas Localization and Segmentation
- Learning Bidirectional LSTM Networks for Synthesizing 3D Mesh Animation Sequences
- Unsupervised Learning of 3D Point Set Registration
- Pointwise Convolutional Neural Networks
- A 4D Light-Field Dataset and CNN Architectures for Material Recognition
- NormalNet: Learning-based Normal Filtering for Mesh Denoising
- Point2Node: Correlation Learning of Dynamic-Node for Point Cloud Feature Modeling
- TearingNet: Point Cloud Autoencoder to Learn Topology-Friendly Representations
- End-to-End Multi-View Networks for Text Classification
- A Graph-CNN for 3D Point Cloud Classification
- MeshNet: Mesh Neural Network for 3D Shape Representation
- Cross-modal Subspace Learning for Fine-grained Sketch-based Image Retrieval
- Zero in on Shape: A Generic 2D-3D Instance Similarity Metric learned from Synthetic Data
- 3DMV: Joint 3D-Multi-View Prediction for 3D Semantic Scene Segmentation
- Canonical and Compact Point Cloud Representation for Shape Classification
- Efficient Globally Optimal 2D-to-3D Deformable Shape Matching
- Mesh-based Autoencoders for Localized Deformation Component Analysis
- PVRNet: Point-View Relation Neural Network for 3D Shape Recognition
- 3DContextNet: K-d Tree Guided Hierarchical Learning of Point Clouds Using Local and Global Contextual Cues
- 3D Object Classification via Spherical Projections
- Angular Triplet-Center Loss for Multi-view 3D Shape Retrieval
- Learning a Hierarchical Latent-Variable Model of 3D Shapes
- Learning Local Shape Descriptors from Part Correspondences With Multi-view Convolutional Networks
- What can we learn about CNNs from a large scale controlled object dataset?
- Multi-view Laplacian Eigenmaps Based on Bag-of-Neighbors For RGBD Human Emotion Recognition
- Deep Cross-modality Adaptation via Semantics Preserving Adversarial Learning for Sketch-based 3D Shape Retrieval
- Multi-Kernel Diffusion CNNs for Graph-Based Learning on Point Clouds
- Automated X-ray Image Analysis for Cargo Security: Critical Review and Future Promise
- Beam Search for Learning a Deep Convolutional Neural Network of 3D Shapes
- Novel Perception Algorithmic Framework For Object Identification and Tracking In Autonomous Navigation
- Local-Area-Learning Network: Meaningful Local Areas for Efficient Point Cloud Analysis
- The importance of silhouette optimization in 3D shape reconstruction system from multiple object scenes
- cvpaper.challenge in 2016: Futuristic Computer Vision through 1,600 Papers Survey
- Multi-View Product Image Search Using Deep ConvNets Representations
- Shape-Oriented Convolution Neural Network for Point Cloud Analysis
- 3D Topology Transformation with Generative Adversarial Networks
- Zero-shot Learning of 3D Point Cloud Objects
- NeuroView: Explainable Deep Network Decision Making
- 3D Shape Retrieval via Irrelevance Filtering and Similarity Ranking (IF/SR)
- Human Recognition Using Face in Computed Tomography
- Exploit Clues from Views: Self-Supervised and Regularized Learning for Multiview Object Recognition