Publications (50)
Snapmoji: Instant Generation of Animatable Dual-Stylized Avatars
Eric M. Chen, Di Liu, Sizhuo Ma +8
Despite the increasing popularity of avatar systems such as Snapchat Bitmojis, existing production avatar platforms face several limitations, such as a limited number of predefined…
History-Guided Video Diffusion
Kiwhan Song, Boyuan Chen, Max Simchowitz +3
Classifier-free guidance (CFG) is a key technique for improving conditional generation in diffusion models, enabling more accurate control while enhancing sample quality. It is nat…
Intrinsic Image Diffusion for Indoor Single-view Material Estimation
Peter Kocsis, Vincent Sitzmann, Matthias NieÃner
We present Intrinsic Image Diffusion, a generative model for appearance decomposition of indoor scenes. Given a single input view, we sample multiple possible material explanations…
Dirty Pixels: Towards End-to-End Image Processing and Perception
Steven Diamond, Vincent Sitzmann, Frank Julca-Aguilar +3
Real-world imaging systems acquire measurements that are degraded by noise, optical aberrations, and other imperfections that make image processing for human viewing and higher-lev…
MetaSDF: Meta-learning Signed Distance Functions
Vincent Sitzmann, Eric R. Chan, Richard Tucker +2
Neural implicit shape representations are an emerging paradigm that offers many potential benefits over conventional discrete representations, including memory efficiency at a high…
Selective Underfitting in Diffusion Models
Kiwhan Song, Jaeyeon Kim, Sitan Chen +3
Diffusion models have emerged as the principal paradigm for generative modeling across various domains. During training, they learn the score function, which in turn is used to gen…
Implicit Neural Representations with Periodic Activation Functions
Vincent Sitzmann, Julien N. P. Martel, Alexander W. Bergman +2
Implicitly defined, continuous, differentiable signal representations parameterized by neural networks have emerged as a powerful paradigm, offering many possible benefits over con…
Scaling View Synthesis Transformers
Evan Kim, Hyunwoo Ryu, Thomas W. Mitchel +1
Geometry-free view synthesis transformers have recently achieved state-of-the-art performance in Novel View Synthesis (NVS), outperforming traditional approaches that rely on expli…
How do people explore virtual environments?
Vincent Sitzmann, Ana Serrano, Amy Pavel +4
Understanding how people explore immersive virtual environments is crucial for many applications, such as designing virtual reality (VR) content, developing new compression algorit…
Meschers: Geometry Processing of Impossible Objects
Ana Dodik, Isabella Yu, Kartik Chandra +4
Impossible objects, geometric constructions that humans can perceive but that cannot exist in real life, have been a topic of intrigue in visual arts, perception, and graphics, yet…
Neural Isometries: Taming Transformations for Equivariant ML
Thomas W. Mitchel, Michael Taylor, Vincent Sitzmann
Real-world geometry and 3D vision tasks are replete with challenging symmetries that defy tractable analytical expression. In this paper, we introduce Neural Isometries, an autoenc…
Score Distillation via Reparametrized DDIM
Artem Lukoianov, Haitz Sáez de Ocáriz Borde, Kristjan Greenewald +4
While 2D diffusion models generate realistic, high-detail images, 3D shape generation methods like Score Distillation Sampling (SDS) built on these 2D diffusion models produce cart…
Meshtryoshka: Differentiable Rendering of Real-World Scenes via Mesh Rasterization
David Charatan, Daniel Xu, Richard Szeliski +2
Differentiable rendering has emerged as a powerful approach for 3D reconstruction and novel view synthesis. State-of-the-art differentiable rendering methods combine a variety of c…
Turning Video Models into Generalist Robot Policies
Sizhe Lester Li, Evan Kim, Xingjian Bai +4
Video generative models have emerged as a promising robotics backbone, capable of generating videos that depict the completion of complex tasks across embodiments and environments.…
Diffusion Forcing: Next-token Prediction Meets Full-Sequence Diffusion
Boyuan Chen, Diego Marti Monso, Yilun Du +3
This paper presents Diffusion Forcing, a new training paradigm where a diffusion model is trained to denoise a set of tokens with independent per-token noise levels. We apply Diffu…
State of the Art on Neural Rendering
Ayush Tewari, Ohad Fried, Justus Thies +16
Efficient rendering of photo-realistic virtual worlds is a long standing effort of computer graphics. Modern graphics techniques have succeeded in synthesizing photo-realistic imag…
Large Video Planner Enables Generalizable Robot Control
Boyuan Chen, Tianyuan Zhang, Haoran Geng +9
General-purpose robots require decision-making models that generalize across diverse tasks and environments. Recent works build robot foundation models by extending multimodal larg…
pixelSplat: 3D Gaussian Splats from Image Pairs for Scalable Generalizable 3D Reconstruction
David Charatan, Sizhe Li, Andrea Tagliasacchi +1
We introduce pixelSplat, a feed-forward model that learns to reconstruct 3D radiance fields parameterized by 3D Gaussian primitives from pairs of images. Our model features real-ti…
MilliVid: Hierarchical Latents for Long-Range Consistency in Video Generation
Ishaan Preetam Chandratreya, David Charatan, Basile Van Hoorick +4
Video generative models have become increasingly powerful, but long-range consistency remains challenging to achieve because even a few dozen frames require impractically long tran…
Kubric: A scalable dataset generator
Klaus Greff, Francois Belletti, Lucas Beyer +32
Data is the driving force of machine learning, with the amount and quality of training data often being more important for the performance of a system than architecture and trainin…
Dataset Distillation for Pre-Trained Self-Supervised Vision Models
George Cazenavette, Antonio Torralba, Vincent Sitzmann
The task of dataset distillation aims to find a small set of synthetic images such that training a model on them reproduces the performance of the same model trained on a much larg…
Controlling diverse robots by inferring Jacobian fields with deep networks
Sizhe Lester Li, Annan Zhang, Boyuan Chen +4
Mirroring the complex structures and diverse functions of natural organisms is a long-standing challenge in robotics. Modern fabrication techniques have greatly expanded the feasib…
DeLiRa: Self-Supervised Depth, Light, and Radiance Fields
Vitor Guizilini, Igor Vasiljevic, Jiading Fang +4
Differentiable volumetric rendering is a powerful paradigm for 3D reconstruction and novel view synthesis. However, standard volume rendering approaches struggle with degenerate ge…
Generative View Stitching
Chonghyuk Song, Michal Stary, Boyuan Chen +2
Autoregressive video diffusion models are capable of long rollouts that are stable and consistent with history, but they are unable to guide the current generation with conditionin…
Light Field Networks: Neural Scene Representations with Single-Evaluation Rendering
Vincent Sitzmann, Semon Rezchikov, William T. Freeman +2
Inferring representations of 3D scenes from 2D observations is a fundamental problem of computer graphics, computer vision, and artificial intelligence. Emerging 3D-structured neur…
DittoGym: Learning to Control Soft Shape-Shifting Robots
Suning Huang, Boyuan Chen, Huazhe Xu +1
Robot co-design, where the morphology of a robot is optimized jointly with a learned policy to solve a specific task, is an emerging area of research. It holds particular promise f…
Neural Fields in Visual Computing and Beyond
Yiheng Xie, Towaki Takikawa, Shunsuke Saito +7
Recent advances in machine learning have created increasing interest in solving visual computing problems using a class of coordinate-based neural networks that parametrize physica…
Learning Signal-Agnostic Manifolds of Neural Fields
Yilun Du, Katherine M. Collins, Joshua B. Tenenbaum +1
Deep neural networks have been used widely to learn the latent structure of datasets, across modalities such as images, shapes, and audio signals. However, existing models are gene…
Locality in Image Diffusion Models Emerges from Data Statistics
Artem Lukoianov, Chenyang Yuan, Justin Solomon +1
Recent work has shown that the generalization ability of image diffusion models arises from the locality properties of the trained neural network. In particular, when denoising a p…
Neural Groundplans: Persistent Neural Scene Representations from a Single Image
Prafull Sharma, Ayush Tewari, Yilun Du +7
We present a method to map 2D image observations of a scene to a persistent 3D scene representation, enabling novel view synthesis and disentangled representation of the movable an…
Learning to Render Novel Views from Wide-Baseline Stereo Pairs
Yilun Du, Cameron Smith, Ayush Tewari +1
We introduce a method for novel view synthesis given only a single wide-baseline stereo image pair. In this challenging regime, 3D scene points are regularly observed only once, re…
Approaching human 3D shape perception with neurally mappable models
Thomas P. O'Connell, Tyler Bonnen, Yoni Friedman +4
Humans effortlessly infer the 3D shape of objects. What computations underlie this ability? Although various computational models have been proposed, none of them capture the human…
Unrolled Optimization with Deep Priors
Steven Diamond, Vincent Sitzmann, Felix Heide +1
A broad class of problems at the core of computational imaging, sensing, and low-level computer vision reduces to the inverse problem of extracting latent images that follow a prio…
Robust Biharmonic Skinning Using Geometric Fields
Ana Dodik, Vincent Sitzmann, Justin Solomon +1
Bounded bihramonic weights are a popular tool used to rig and deform characters for animation, to compute reduced-order simulations, and to define feature descriptors for geometry…
Semantic Implicit Neural Scene Representations With Semi-Supervised Training
Amit Kohli, Vincent Sitzmann, Gordon Wetzstein
The recent success of implicit neural scene representations has presented a viable new method for how we capture and store 3D scenes. Unlike conventional 3D representations, such a…
Understanding Multi-View Transformers
Michal Stary, Julien Gaubil, Ayush Tewari +1
Multi-view transformers such as DUSt3R are revolutionizing 3D vision by solving 3D tasks in a feed-forward manner. However, contrary to previous optimization-based pipelines, the i…
Variational Barycentric Coordinates
Ana Dodik, Oded Stein, Vincent Sitzmann +1
We propose a variational technique to optimize for generalized barycentric coordinates that offers additional control compared to existing models. Prior work represents barycentric…
Unsupervised Discovery and Composition of Object Light Fields
Cameron Smith, Hong-Xing Yu, Sergey Zakharov +4
Neural scene representations, both continuous and discrete, have recently emerged as a powerful new paradigm for 3D scene understanding. Recent efforts have tackled unsupervised di…
Advances in Neural Rendering
Ayush Tewari, Justus Thies, Ben Mildenhall +14
Synthesizing photo-realistic images and videos is at the heart of computer graphics and has been the focus of decades of research. Traditionally, synthetic images of a scene are ge…
FlowMap: High-Quality Camera Poses, Intrinsics, and Depth via Gradient Descent
Cameron Smith, David Charatan, Ayush Tewari +1
This paper introduces FlowMap, an end-to-end differentiable method that solves for precise camera poses, camera intrinsics, and per-frame dense depth of a video sequence. Our metho…
Scene Representation Networks: Continuous 3D-Structure-Aware Neural Scene Representations
Vincent Sitzmann, Michael Zollhöfer, Gordon Wetzstein
Unsupervised learning with generative models has the potential of discovering rich representations of 3D scenes. While geometric deep learning has explored 3D-structure-aware repre…
Neural Descriptor Fields: SE(3)-Equivariant Object Representations for Manipulation
Anthony Simeonov, Yilun Du, Andrea Tagliasacchi +4
We present Neural Descriptor Fields (NDFs), an object representation that encodes both points and relative poses between an object and a target (such as a robot gripper or a rack u…
True Self-Supervised Novel View Synthesis is Transferable
Thomas W. Mitchel, Hyunwoo Ryu, Vincent Sitzmann
In this paper, we identify that the key criterion for determining whether a model is truly capable of novel view synthesis (NVS) is transferability: Whether any pose representation…
Diffusion with Forward Models: Solving Stochastic Inverse Problems Without Direct Supervision
Ayush Tewari, Tianwei Yin, George Cazenavette +5
Denoising diffusion models are a powerful type of generative models used to capture complex distributions of real-world signals. However, their applicability is limited to scenario…
Decomposing NeRF for Editing via Feature Field Distillation
Sosuke Kobayashi, Eiichi Matsumoto, Vincent Sitzmann
Emerging neural radiance fields (NeRF) are a promising scene representation for computer graphics, enabling high-quality 3D reconstruction and novel view synthesis from image obser…
3D Neural Scene Representations for Visuomotor Control
Yunzhu Li, Shuang Li, Vincent Sitzmann +2
Humans have a strong intuitive understanding of the 3D environment around us. The mental model of the physics in our brain applies to objects of different materials and enables us…
DeepVoxels: Learning Persistent 3D Feature Embeddings
Vincent Sitzmann, Justus Thies, Felix Heide +3
In this work, we address the lack of 3D understanding of generative neural networks by introducing a persistent 3D feature embedding for view synthesis. To this end, we propose Dee…
Deep Medial Fields
Daniel Rebain, Ke Li, Vincent Sitzmann +3
Implicit representations of geometry, such as occupancy fields or signed distance fields (SDF), have recently re-gained popularity in encoding 3D solid shape in a functional form.…
Movie Editing and Cognitive Event Segmentation in Virtual Reality Video
Ana Serrano, Vincent Sitzmann, Jaime Ruiz-Borau +3
Traditional cinematography has relied for over a century on a well-established set of editing rules, called continuity editing, to create a sense of situational continuity. Despite…
FlowCam: Training Generalizable 3D Radiance Fields without Camera Poses via Pixel-Aligned Scene Flow
Cameron Smith, Yilun Du, Ayush Tewari +1
Reconstruction of 3D neural fields from posed images has emerged as a promising method for self-supervised representation learning. The key challenge preventing the deployment of t…