Learning to Reconstruct and Segment 3D Objects
arXiv:2010.09582
Abstract
To endow machines with the ability to perceive the real-world in a three dimensional representation as we do as humans is a fundamental and long-standing topic in Artificial Intelligence. Given different types of visual inputs such as images or point clouds acquired by 2D/3D sensors, one important goal is to understand the geometric structure and semantics of the 3D environment. Traditional approaches usually leverage hand-crafted features to estimate the shape and semantics of objects or scenes. However, they are difficult to generalize to novel objects and scenarios, and struggle to overcome critical issues caused by visual occlusions. By contrast, we aim to understand scenes and the objects within them by learning general and robust representations using deep neural networks, trained on large-scale real-world 3D data. To achieve these aims, this thesis makes three core contributions from object-level 3D shape estimation from single or multiple views to scene-level semantic understanding.
DPhil (PhD) Thesis 2020, University of Oxford https://ora.ox.ac.uk/objects/uuid:5f9cd30d-0ee7-412d-ba49-44f5fd76bf28
References in corpus (15)
- Conditional Generative Adversarial Nets
- Learning a Probabilistic Latent Space of Object Shapes via 3D Generative-Adversarial Modeling
- Deep Convolutional Inverse Graphics Network
- MarrNet: 3D Shape Reconstruction via 2.5D Sketches
- Learning What and Where to Draw
- Learning Efficient Point Cloud Generation for Dense 3D Object Reconstruction
- SceneNet RGB-D: 5M Photorealistic Images of Synthetic Indoor Trajectories with Ground Truth
- Learning to Infer Implicit Surfaces without 3D Supervision
- MASC: Multi-scale Affinity with Sparse Convolution for 3D Instance Segmentation
- Stochastic Optimization of Sorting Networks via Continuous Relaxations
- Exploiting Local and Global Structure for Point Cloud Semantic Segmentation with Contextual Point Representations
- SemanticPOSS: A Point Cloud Dataset with Large Quantity of Dynamic Instances
- Toward 3D Object Reconstruction from Stereo Images
- AdvectiveNet: An Eulerian-Lagrangian Fluidic reservoir for Point Cloud Processing
- Learning and Memorizing Representative Prototypes for 3D Point Cloud Semantic and Instance Segmentation