OctNet: Learning Deep 3D Representations at High Resolutions
arXiv:1611.05009
Abstract
We present OctNet, a representation for deep learning with sparse 3D data. In contrast to existing models, our representation enables 3D convolutional networks which are both deep and high resolution. Towards this goal, we exploit the sparsity in the input data to hierarchically partition the space using a set of unbalanced octrees where each leaf node stores a pooled feature representation. This allows to focus memory allocation and computation to the relevant dense regions and enables deeper networks without compromising resolution. We demonstrate the utility of our OctNet representation by analyzing the impact of resolution on several 3D tasks including 3D object classification, orientation estimation and point cloud labeling.
CVPR 2017 camera ready
References in corpus (6)
- Deep Residual Learning for Image Recognition
- Generative and Discriminative Voxel Modeling with Convolutional Neural Networks
- Spatially-sparse convolutional neural networks
- A Large Dataset of Object Scans
- 3D ShapeNets: A Deep Representation for Volumetric Shapes
- Generic 3D Representation via Pose Estimation and Matching
Cited by in corpus (20)
- PointNet++: Deep Hierarchical Feature Learning on Point Sets in a Metric Space
- SPLATNet: Sparse Lattice Networks for Point Cloud Processing
- Pix3D: Dataset and Methods for Single-Image 3D Shape Modeling
- Mining Point Cloud Local Structures by Kernel Correlation and Graph Pooling
- Hierarchical Surface Prediction for 3D Object Reconstruction
- OctNetFusion: Learning Depth Fusion from Data
- Compact Model Representation for 3D Reconstruction
- Octree Generating Networks: Efficient Convolutional Architectures for High-resolution 3D Outputs
- DIST: Rendering Deep Implicit Signed Distance Function with Differentiable Sphere Tracing
- Multi-View Deep Learning for Consistent Semantic Mapping with RGB-D Cameras
- DAR-Net: Dynamic Aggregation Network for Semantic Scene Segmentation
- NormalNet: Learning-based Normal Filtering for Mesh Denoising
- A Spatial Mapping Algorithm with Applications in Deep Learning-Based Structure Classification
- Local Deep Implicit Functions for 3D Shape
- Shape Generation using Spatially Partitioned Point Clouds
- What can I do here? Leveraging Deep 3D saliency and geometry for fast and scalable multiple affordance detection
- SurfConv: Bridging 3D and 2D Convolution for RGBD Images
- 3D Object Classification via Spherical Projections
- AdvectiveNet: An Eulerian-Lagrangian Fluidic reservoir for Point Cloud Processing
- MARNet: Multi-Abstraction Refinement Network for 3D Point Cloud Analysis