Publications (44)
Lightplane: Highly-Scalable Components for Neural 3D Fields
Ang Cao, Justin Johnson, Andrea Vedaldi +1
Contemporary 3D research, particularly in reconstruction and generation, heavily relies on 2D images for inputs or supervision. However, current designs for these 2D-3D mapping are…
iSDF: Real-Time Neural Signed Distance Fields for Robot Perception
Joseph Ortiz, Alexander Clegg, Jing Dong +4
We present iSDF, a continual learning system for real-time signed distance field (SDF) reconstruction. Given a stream of posed depth images from a moving camera, it trains a random…
Canonical 3D Deformer Maps: Unifying parametric and non-parametric methods for dense weakly-supervised category reconstruction
David Novotny, Roman Shapovalov, Andrea Vedaldi
We propose the Canonical 3D Deformer Map, a new representation of the 3D shape of common object categories that can be learned from a collection of 2D images of independent objects…
Learning 3D Object Categories by Looking Around Them
David Novotny, Diane Larlus, Andrea Vedaldi
Traditional approaches for learning 3D object categories use either synthetic data or manual supervision. In this paper, we propose a method which does not require manual annotatio…
DensePose 3D: Lifting Canonical Surface Maps of Articulated Objects to the Third Dimension
Roman Shapovalov, David Novotny, Benjamin Graham +2
We tackle the problem of monocular 3D reconstruction of articulated objects like humans and animals. We contribute DensePose 3D, a method that can learn such reconstructions in a w…
Real-time volumetric rendering of dynamic humans
Ignacio Rocco, Iurii Makarov, Filippos Kokkinos +4
We present a method for fast 3D reconstruction and real-time rendering of dynamic humans from monocular videos with accompanying parametric body fits. Our method can reconstruct a…
UnCommon Objects in 3D
Xingchen Liu, Piyush Tayal, Jianyuan Wang +10
We introduce Uncommon Objects in 3D (uCO3D), a new object-centric dataset for 3D deep learning and 3D generative AI. uCO3D is the largest publicly-available collection of high-reso…
Unsupervised Learning of 3D Object Categories from Videos in the Wild
Philipp Henzler, Jeremy Reizenstein, Patrick Labatut +4
Our goal is to learn a deep network that, given a small number of images of an object of a given category, reconstructs it in 3D. While several recent works have obtained analogous…
RidgeSfM: Structure from Motion via Robust Pairwise Matching Under Depth Uncertainty
Benjamin Graham, David Novotny
We consider the problem of simultaneously estimating a dense depth map and camera pose for a large set of images of an indoor scene. While classical SfM pipelines rely on a two-ste…
C3DPO: Canonical 3D Pose Networks for Non-Rigid Structure From Motion
David Novotny, Nikhila Ravi, Benjamin Graham +2
We propose C3DPO, a method for extracting 3D models of deformable objects from 2D keypoint annotations in unconstrained images. We do so by learning a deep network that reconstruct…
Learning the semantic structure of objects from Web supervision
David Novotny, Diane Larlus, Andrea Vedaldi
While recent research in image understanding has often focused on recognizing more types of objects, understanding more about the objects is just as important. Recognizing object p…
Self-Supervised Correspondence Estimation via Multiview Registration
Mohamed El Banani, Ignacio Rocco, David Novotny +4
Video provides us with the spatio-temporal consistency needed for visual learning. Recent approaches have utilized this signal to learn correspondence estimation from close-by fram…
ActionMesh: Animated 3D Mesh Generation with Temporal 3D Diffusion
Remy Sabathier, David Novotny, Niloy J. Mitra +1
Generating animated 3D objects is at the heart of many applications, yet most advanced works are typically difficult to apply in practice because of their limited setup, their long…
GOEmbed: Gradient Origin Embeddings for Representation Agnostic 3D Feature Learning
Animesh Karnewar, Roman Shapovalov, Tom Monnier +3
Encoding information from 2D views of an object into a 3D representation is crucial for generalized 3D feature extraction. Such features can then enable 3D reconstruction, 3D gener…
WildCAT3D: Appearance-Aware Multi-View Diffusion in the Wild
Morris Alper, David Novotny, Filippos Kokkinos +2
Despite recent advances in sparse novel view synthesis (NVS) applied to object-centric scenes, scene-level NVS remains a challenge. A central issue is the lack of available clean m…
Visual Geometry Grounded Deep Structure From Motion
Jianyuan Wang, Nikita Karaev, Christian Rupprecht +1
Structure-from-motion (SfM) is a long-standing problem in the computer vision community, which aims to reconstruct the camera poses and 3D structure of a scene from a set of uncons…
Common Objects in 3D: Large-Scale Learning and Evaluation of Real-life 3D Category Reconstruction
Jeremy Reizenstein, Roman Shapovalov, Philipp Henzler +3
Traditional approaches for learning 3D object categories have been predominantly trained and evaluated on synthetic datasets due to the unavailability of real 3D-annotated category…
Animal Avatars: Reconstructing Animatable 3D Animals from Casual Videos
Remy Sabathier, Niloy J. Mitra, David Novotny
We present a method to build animatable dog avatars from monocular videos. This is challenging as animals display a range of (unpredictable) non-rigid movements and have a variety…
NeuroMorph: Unsupervised Shape Interpolation and Correspondence in One Go
Marvin Eisenberger, David Novotny, Gael Kerchenbaum +4
We present NeuroMorph, a new neural network architecture that takes as input two 3D shapes and produces in one go, i.e. in a single feed forward pass, a smooth interpolation and po…
Continuous Surface Embeddings
Natalia Neverova, David Novotny, Vasil Khalidov +3
In this work, we focus on the task of learning and representing dense correspondences in deformable object categories. While this problem has been considered before, solutions so f…
ViewDiff: 3D-Consistent Image Generation with Text-to-Image Models
Lukas Höllein, Aljaž BožiÄ, Norman Müller +5
3D asset generation is getting massive amounts of attention, inspired by the recent success of text-guided 2D content creation. Existing text-to-3D methods use pretrained text-to-i…
SceneScribe-1M: A Large-Scale Video Dataset with Comprehensive Geometric and Semantic Annotations
Yunnan Wang, Kecheng Zheng, Jianyuan Wang +8
The convergence of 3D geometric perception and video synthesis has created an unprecedented demand for large-scale video data that is rich in both semantic and spatio-temporal info…
HoloFusion: Towards Photo-realistic 3D Generative Modeling
Animesh Karnewar, Niloy J. Mitra, Andrea Vedaldi +1
Diffusion-based image generators can now produce high-quality and diverse samples, but their success has yet to fully translate to 3D generation: existing diffusion methods can eit…
Meta 3D Gen
Raphael Bensadoun, Tom Monnier, Yanir Kleiman +17
We introduce Meta 3D Gen (3DGen), a new state-of-the-art, fast pipeline for text-to-3D asset generation. 3DGen offers 3D asset creation with high prompt fidelity and high-quality 3…
LIM: Large Interpolator Model for Dynamic Reconstruction
Remy Sabathier, Niloy J. Mitra, David Novotny
Reconstructing dynamic assets from video data is central to many in computer vision and graphics tasks. Existing 4D reconstruction approaches are limited by category-specific model…
Detecting and Receiving Phase Modulated Signals with a Rydberg Atom-Based Mixer
Christopher L. Holloway, Matthew T. Simons, Joshua A. Gordon +1
Recently, we introduced a Rydberg-atom based mixer capable of detecting and measuring the phase of a radio-frequency field through the electromagnetically induced transparency (EIT…
VGGT: Visual Geometry Grounded Transformer
Jianyuan Wang, Minghao Chen, Nikita Karaev +3
We present VGGT, a feed-forward neural network that directly infers all key 3D attributes of a scene, including camera parameters, point maps, depth maps, and 3D point tracks, from…
VGGT-
Jianyuan Wang, Minghao Chen, Shangzhan Zhang +7
Recent feed-forward reconstruction models, such as VGGT, have proven competitive with traditional optimization-based reconstructors while also providing geometry-aware features use…
PartGen: Part-level 3D Generation and Reconstruction with Multi-View Diffusion Models
Minghao Chen, Roman Shapovalov, Iro Laina +4
Text- or image-to-3D generators and 3D scanners can now produce 3D assets with high-quality shapes and textures. These assets typically consist of a single, fused representation, l…
Augmenting Implicit Neural Shape Representations with Explicit Deformation Fields
Matan Atzmon, David Novotny, Andrea Vedaldi +1
Implicit neural representation is a recent approach to learn shape collections as zero level-sets of neural networks, where each shape is represented by a latent code. So far, the…
Meta 3D AssetGen: Text-to-Mesh Generation with High-Quality Geometry, Texture, and PBR Materials
Yawar Siddiqui, Tom Monnier, Filippos Kokkinos +8
We present Meta 3D AssetGen (AssetGen), a significant advancement in text-to-3D generation which produces faithful, high-quality meshes with texture and material control. Compared…
Discovering Relationships between Object Categories via Universal Canonical Maps
Natalia Neverova, Artsiom Sanakoyeu, Patrick Labatut +2
We tackle the problem of learning the geometry of multiple categories of deformable objects jointly. Recent work has shown that it is possible to learn a unified dense pose predict…
AnchorNet: A Weakly Supervised Network to Learn Geometry-sensitive Features For Semantic Matching
David Novotny, Diane Larlus, Andrea Vedaldi
Despite significant progress of deep learning in recent years, state-of-the-art semantic matching methods still rely on legacy features such as SIFT or HoG. We argue that the stron…
Nerfels: Renderable Neural Codes for Improved Camera Pose Estimation
Gil Avraham, Julian Straub, Tianwei Shen +7
This paper presents a framework that combines traditional keypoint-based camera pose optimization with an invertible neural rendering mechanism. Our proposed 3D scene representatio…
PoseDiffusion: Solving Pose Estimation via Diffusion-aided Bundle Adjustment
Jianyuan Wang, Christian Rupprecht, David Novotny
Camera pose estimation is a long-standing computer vision problem that to date often relies on classical methods, such as handcrafted keypoint matching, RANSAC and bundle adjustmen…
Replay: Multi-modal Multi-view Acted Videos for Casual Holography
Roman Shapovalov, Yanir Kleiman, Ignacio Rocco +6
We introduce Replay, a collection of multi-view, multi-modal videos of humans interacting socially. Each scene is filmed in high production quality, from different viewpoints with…
Twinner: Shining Light on Digital Twins in a Few Snaps
Jesus Zarzar, Tom Monnier, Roman Shapovalov +2
We present the first large reconstruction model, Twinner, capable of recovering a scene's illumination as well as an object's geometry and material properties from only a few posed…
HoloDiffusion: Training a 3D Diffusion Model using 2D Images
Animesh Karnewar, Andrea Vedaldi, David Novotny +1
Diffusion models have emerged as the best approach for generative modeling of 2D images. Part of their success is due to the possibility of training them on millions if not billion…
Cascaded Sparse Spatial Bins for Efficient and Effective Generic Object Detection
David Novotny, Jiri Matas
A novel efficient method for extraction of object proposals is introduced. Its "objectness" function exploits deep spatial pyramid features, a novel fast-to-compute HoG-based edge…
Common Pets in 3D: Dynamic New-View Synthesis of Real-Life Deformable Categories
Samarth Sinha, Roman Shapovalov, Jeremy Reizenstein +4
Obtaining photorealistic reconstructions of objects from sparse views is inherently ambiguous and can only be achieved by learning suitable reconstruction priors. Earlier works on…
Self-supervised Learning of Geometrically Stable Features Through Probabilistic Introspection
David Novotny, Samuel Albanie, Diane Larlus +1
Self-supervision can dramatically cut back the amount of manually-labelled data required to train deep neural networks. While self-supervision has usually been considered for tasks…
Accelerating 3D Deep Learning with PyTorch3D
Nikhila Ravi, Jeremy Reizenstein, David Novotny +4
Deep learning has significantly improved 2D image recognition. Extending into 3D may advance many new applications including autonomous vehicles, virtual and augmented reality, aut…
Semi-convolutional Operators for Instance Segmentation
David Novotny, Samuel Albanie, Diane Larlus +1
Object detection and instance segmentation are dominated by region-based methods such as Mask RCNN. However, there is a growing interest in reducing these problems to pixel labelin…
3D Multi-bodies: Fitting Sets of Plausible 3D Human Models to Ambiguous Image Data
Benjamin Biggs, Sébastien Ehrhadt, Hanbyul Joo +3
We consider the problem of obtaining dense 3D reconstructions of humans from single and partially occluded views. In such cases, the visual evidence is usually insufficient to iden…