Implicit Mesh Reconstruction from Unannotated Image Collections
arXiv:2007.08504
Abstract
We present an approach to infer the 3D shape, texture, and camera pose for an object from a single RGB image, using only category-level image collections with foreground masks as supervision. We represent the shape as an image-conditioned implicit function that transforms the surface of a sphere to that of the predicted mesh, while additionally predicting the corresponding texture. To derive supervisory signal for learning, we enforce that: a) our predictions when rendered should explain the available image evidence, and b) the inferred 3D structure should be geometrically consistent with learned pixel to surface mappings. We empirically show that our approach improves over prior work that leverages similar supervision, and in fact performs competitively to methods that use stronger supervision. Finally, as our method enables learning with limited supervision, we qualitatively demonstrate its applicability over a set of about 30 object categories.
Project page: https://shubhtuls.github.io/imr/
References in corpus (1)
Cited by in corpus (5)
- NeRS: Neural Reflectance Surfaces for Sparse-view 3D Reconstruction in the Wild
- Learning 3D Dense Correspondence via Canonical Point Autoencoder
- StrobeNet: Category-Level Multiview Reconstruction of Articulated Objects
- Do 2D GANs Know 3D Shape? Unsupervised 3D shape reconstruction from 2D Image GANs
- Birds of a Feather: Capturing Avian Shape Models from Images