SeqXY2SeqZ: Structure Learning for 3D Shapes by Sequentially Predicting 1D Occupancy Segments From 2D Coordinates
arXiv:2003.05559
Abstract
Structure learning for 3D shapes is vital for 3D computer vision. State-of-the-art methods show promising results by representing shapes using implicit functions in 3D that are learned using discriminative neural networks. However, learning implicit functions requires dense and irregular sampling in 3D space, which also makes the sampling methods affect the accuracy of shape reconstruction during test. To avoid dense and irregular sampling in 3D, we propose to represent shapes using 2D functions, where the output of the function at each 2D location is a sequence of line segments inside the shape. Our approach leverages the power of functional representations, but without the disadvantage of 3D sampling. Specifically, we use a voxel tubelization to represent a voxel grid as a set of tubes along any one of the X, Y, or Z axes. Each tube can be indexed by its 2D coordinates on the plane spanned by the other two axes. We further simplify each tube into a sequence of occupancy segments. Each occupancy segment consists of successive voxels occupied by the shape, which leads to a simple representation of its 1D start and end location. Given the 2D coordinates of the tube and a shape feature as condition, this representation enables us to learn 3D shape structures by sequentially predicting the start and end locations of each occupancy segment in the tube. We implement this approach using a Seq2Seq model with attention, called SeqXY2SeqZ, which learns the mapping from a sequence of 2D coordinates along two arbitrary axes to a sequence of 1D locations along the third axis. SeqXY2SeqZ not only benefits from the regularity of voxel grids in training and testing, but also achieves high memory efficiency. Our experiments show that SeqXY2SeqZ outperforms the state-ofthe-art methods under widely used benchmarks.
References in corpus (16)
- Learning a Probabilistic Latent Space of Object Shapes via 3D Generative-Adversarial Modeling
- Perspective Transformer Nets: Learning Single-View 3D Object Reconstruction without 3D Supervision
- Learning to Predict 3D Objects with an Interpolation-based Differentiable Renderer
- Differentiable Surface Splatting for Point-based Geometry Processing
- MarrNet: 3D Shape Reconstruction via 2.5D Sketches
- Learning Efficient Point Cloud Generation for Dense 3D Object Reconstruction
- Soft Rasterizer: A Differentiable Renderer for Image-based 3D Reasoning
- Learning to Infer Implicit Surfaces without 3D Supervision
- Learning to Reconstruct Shapes from Unseen Classes
- Soft Rasterizer: Differentiable Rendering for Unsupervised Single-View Mesh Reconstruction
- Matryoshka Networks: Predicting 3D Geometry via Nested Shape Layers
- Deep Level Sets: Implicit Surface Representations for 3D Shape Inference
- 3DN: 3D Deformation Network
- Multi-Angle Point Cloud-VAE: Unsupervised Feature Learning for 3D Point Clouds from Multiple Angles by Joint Self-Reconstruction and Half-to-Half Prediction
- 3DViewGraph: Learning Global Features for 3D Shapes from A Graph of Unordered Views with Attention
- Parts4Feature: Learning 3D Global Features from Generally Semantic Parts in Multiple Views