papers

Publications (179)

cs.CV2026

World from Motion: Generative Dynamic Gaussian Reconstruction from Monocular Video

Liyuan Zhu, Shengyu Huang, Amrita Mazumdar +6

We present World from Motion, a method for generating freely renderable dynamic 3D Gaussian representations from monocular videos. Our approach conditions a video model on dense, p…

cs.CV2026

Spectral Progressive Diffusion for Efficient Image and Video Generation

Howard Xiao, Brian Chao, Lior Yariv +1

Diffusion models have been shown to implicitly generate visual content autoregressively in the frequency domain, where low-frequency components are generated earlier in the denoisi…

eess.IV2025

Patch-Based Diffusion for Data-Efficient, Radiologist-Preferred MRI Reconstruction

Rohan Sanda, Asad Aali, Andrew Johnston +3

Magnetic resonance imaging (MRI) requires long acquisition times, raising costs, reducing accessibility, and making scans more susceptible to motion artifacts. Diffusion probabilis…

cs.CV2025

Taming Flow-based I2V Models for Creative Video Editing

Xianghao Kong, Hansheng Chen, Yuwei Guo +4

Although image editing techniques have advanced significantly, video editing, which aims to manipulate videos according to user intent, remains an emerging challenge. Most existing…

cs.LG2026

pi-Flow: Policy-Based Few-Step Generation via Imitation Distillation

Hansheng Chen, Kai Zhang, Hao Tan +3

Few-step diffusion or flow-based generative models typically distill a velocity-predicting teacher into a student that predicts a shortcut towards denoised data. This format mismat…

cs.CV2025

Interspatial Attention for Efficient 4D Human Video Generation

Ruizhi Shao, Yinghao Xu, Yujun Shen +5

Generating photorealistic videos of digital humans in a controllable manner is crucial for a plethora of applications. Existing approaches either build on methods that employ templ…

cs.CV2026

Effective Multi-sensor Conditioning for Street-view Novel-view Synthesis

Zhengfei Kuang, Adam Sun, Liyuan Zhu +7

Modern vehicle platforms are equipped with a rich sensor suite, including LiDAR, calibrated multi-camera rigs, and accurate ego-motion, that in principle offers strong signal for r…

cs.CV2023

3D-Aware Video Generation

Sherwin Bahmani, Jeong Joon Park, Despoina Paschalidou +5

Generative models have emerged as an essential building block for many image synthesis and editing tasks. Recent advances in this field have also enabled high-quality 3D or video c…

cs.CV2025

Envision: Embodied Visual Planning via Goal-Imagery Video Diffusion

Yuming Gu, Yizhi Wang, Yining Hong +9

Embodied visual planning aims to enable manipulation tasks by imagining how a scene evolves toward a desired goal and using the imagined trajectories to guide actions. Video diffus…

cs.CV2023

LumiGAN: Unconditional Generation of Relightable 3D Human Faces

Boyang Deng, Yifan Wang, Gordon Wetzstein

Unsupervised learning of 3D human faces from unstructured 2D image data is an active research area. While recent works have achieved an impressive level of photorealism, they commo…

cs.LG2026

Modeling Atomic Conformational Ensembles of Proteins via Test-Time Supervision of Boltz-2 on Cryo-EM Density Maps

Jay Shenoy, Miro Astore, Axel Levy +3

Knowledge of a protein's atomic conformational ensemble is critical to determining its function, yet state-of-the-art ensemble prediction models are limited by lack of high-quality…

cs.CV2024

Generic 3D Diffusion Adapter Using Controlled Multi-View Editing

Hansheng Chen, Ruoxi Shi, Yulin Liu +5

Open-domain 3D object synthesis has been lagging behind image synthesis due to limited data and higher computational complexity. To bridge this gap, recent works have investigated…

cs.CV2021

AutoInt: Automatic Integration for Fast Neural Volume Rendering

David B. Lindell, Julien N. P. Martel, Gordon Wetzstein

Numerical integration is a foundational technique in scientific computing and is at the core of many computer vision applications. Among these applications, neural volume rendering…

cs.GR2020

Gaze-Contingent Ocular Parallax Rendering for Virtual Reality

Robert Konrad, Anastasios Angelopoulos, Gordon Wetzstein

Immersive computer graphics systems strive to generate perceptually realistic user experiences. Current-generation virtual reality (VR) displays are successful in accurately render…

cs.CV2025

CL-Splats: Continual Learning of Gaussian Splatting with Local Optimization

Jan Ackermann, Jonas Kulhanek, Shengqu Cai +5

In dynamic 3D environments, accurately updating scene representations over time is crucial for applications in robotics, mixed reality, and embodied AI. As scenes evolve, efficient…

cs.CV2022

Amortized Inference for Heterogeneous Reconstruction in Cryo-EM

Axel Levy, Gordon Wetzstein, Julien Martel +2

Cryo-electron microscopy (cryo-EM) is an imaging modality that provides unique insights into the dynamics of proteins and other building blocks of life. The algorithmic challenge o…

cs.CV2024

3DitScene: Editing Any Scene via Language-guided Disentangled Gaussian Splatting

Qihang Zhang, Yinghao Xu, Chaoyang Wang +4

Scene image editing is crucial for entertainment, photography, and advertising design. Existing methods solely focus on either 2D individual object or 3D global scene editing. This…

cs.CV2024

GPT-4V(ision) is a Human-Aligned Evaluator for Text-to-3D Generation

Tong Wu, Guandao Yang, Zhibing Li +5

Despite recent advances in text-to-3D generative methods, there is a notable absence of reliable evaluation metrics. Existing metrics usually focus on a single criterion each, such…

cs.CV2025

ByteMorph: Benchmarking Instruction-Guided Image Editing with Non-Rigid Motions

Di Chang, Mingdeng Cao, Yichun Shi +7

Editing images with instructions to reflect non-rigid motions, camera viewpoint shifts, object deformations, human articulations, and complex interactions, poses a challenging yet…

cs.CV2021

pi-GAN: Periodic Implicit Generative Adversarial Networks for 3D-Aware Image Synthesis

Eric R. Chan, Marco Monteiro, Petr Kellnhofer +2

We have witnessed rapid progress on 3D-aware image synthesis, leveraging recent advances in generative visual models and neural rendering. Existing approaches however fall short in…

cs.CV2025

Visual Chronicles: Using Multimodal LLMs to Analyze Massive Collections of Images

Boyang Deng, Songyou Peng, Kyle Genova +4

We present a system using Multimodal LLMs (MLLMs) to analyze a large database with tens of millions of images captured at different times, with the aim of discovering patterns in t…

cs.GR2025

WonderPlay: Dynamic 3D Scene Generation from a Single Image and Actions

Zizhang Li, Hong-Xing Yu, Wei Liu +4

WonderPlay is a novel framework integrating physics simulation with video generation for generating action-conditioned dynamic 3D scenes from a single image. While prior works are…

eess.IV2020

Michelson Holography: Dual-SLM Holography with Camera-in-the-loop Optimization

Suyeon Choi, Jonghyun Kim, Yifan Peng +1

We introduce Michelson Holography (MH), a holographic display technology that optimizes image quality for emerging holographic near-eye displays. Using two spatial light modulators…

cs.CV2024

MegaScenes: Scene-Level View Synthesis at Scale

Joseph Tung, Gene Chou, Ruojin Cai +5

Scene-level novel view synthesis (NVS) is fundamental to many vision and graphics applications. Recently, pose-conditioned diffusion models have led to significant progress by extr…

cs.CV2024

GRM: Large Gaussian Reconstruction Model for Efficient 3D Reconstruction and Generation

Yinghao Xu, Zifan Shi, Wang Yifan +5

We introduce GRM, a large-scale reconstructor capable of recovering a 3D asset from sparse-view images in around 0.1s. GRM is a feed-forward transformer-based model that efficientl…

cs.CV2024

Dynamic Gaussian Marbles for Novel View Synthesis of Casual Monocular Videos

Colton Stearns, Adam Harley, Mikaela Uy +4

Gaussian splatting has become a popular representation for novel-view synthesis, exhibiting clear strengths in efficiency, photometric quality, and compositional edibility. Followi…

cs.CV2021

Keyhole Imaging: Non-Line-of-Sight Imaging and Tracking of Moving Objects Along a Single Optical Path

Christopher A. Metzler, David B. Lindell, Gordon Wetzstein

Non-line-of-sight (NLOS) imaging and tracking is an emerging technology that allows the shape or position of objects around corners or behind diffusers to be recovered from transie…

cs.CV2023

DMV3D: Denoising Multi-View Diffusion using 3D Large Reconstruction Model

Yinghao Xu, Hao Tan, Fujun Luan +8

We propose \textbf{DMV3D}, a novel 3D generation approach that uses a transformer-based 3D large reconstruction model to denoise multi-view diffusion. Our reconstruction model inco…

cs.CV2025

LayerPano3D: Layered 3D Panorama for Hyper-Immersive Scene Generation

Shuai Yang, Jing Tan, Mengchen Zhang +5

3D immersive scene generation is a challenging yet critical task in computer vision and graphics. A desired virtual 3D scene should 1) exhibit omnidirectional view consistency, and…

cs.CV2020

Scene Representation Networks: Continuous 3D-Structure-Aware Neural Scene Representations

Vincent Sitzmann, Michael Zollhöfer, Gordon Wetzstein

Unsupervised learning with generative models has the potential of discovering rich representations of 3D scenes. While geometric deep learning has explored 3D-structure-aware repre…

cs.CV2019

Deep Optics for Monocular Depth Estimation and 3D Object Detection

Julie Chang, Gordon Wetzstein

Depth estimation and 3D object detection are critical for scene understanding but remain challenging to perform with a single image due to the loss of 3D information during image c…

cs.GR2024

Holographic Parallax Improves 3D Perceptual Realism

Dongyeon Kim, Seung-Woo Nam, Suyeon Choi +3

Holographic near-eye displays are a promising technology to solve long-standing challenges in virtual and augmented reality display systems. Over the last few years, many different…

cs.CV2023

DehazeNeRF: Multiple Image Haze Removal and 3D Shape Reconstruction using Neural Radiance Fields

Wei-Ting Chen, Wang Yifan, Sy-Yen Kuo +1

Neural radiance fields (NeRFs) have demonstrated state-of-the-art performance for 3D computer vision tasks, including novel view synthesis and 3D shape reconstruction. However, the…

eess.IV2021

Time-Multiplexed Coded Aperture Imaging: Learned Coded Aperture and Pixel Exposures for Compressive Imaging Systems

Edwin Vargas, Julien N. P. Martel, Gordon Wetzstein +1

Compressive imaging using coded apertures (CA) is a powerful technique that can be used to recover depth, light fields, hyperspectral images and other quantities from a single snap…

cs.ET2025

Roadmap on Neuromorphic Photonics

Daniel Brunner, Bhavin J. Shastri, Mohammed A. Al Qadasi +147

This roadmap consolidates recent advances while exploring emerging applications, reflecting the remarkable diversity of hardware platforms, neuromorphic concepts, and implementatio…

cs.CV2024

Buffer Anytime: Zero-Shot Video Depth and Normal from Image Priors

Zhengfei Kuang, Tianyuan Zhang, Kai Zhang +7

We present Buffer Anytime, a framework for estimation of depth and normal maps (which we call geometric buffers) from video that eliminates the need for paired video--depth and vid…

cs.CV2022

Event Based, Near Eye Gaze Tracking Beyond 10,000Hz

Anastasios N. Angelopoulos, Julien N. P. Martel, Amit P. S. Kohli +2

The cameras in modern gaze-tracking systems suffer from fundamental bandwidth and power limitations, constraining data acquisition speed to 300 Hz realistically. This obstructs the…

cs.CV2025

AIpparel: A Multimodal Foundation Model for Digital Garments

Kiyohiro Nakayama, Jan Ackermann, Timur Levent Kesdogan +6

Apparel is essential to human life, offering protection, mirroring cultural identities, and showcasing personal style. Yet, the creation of garments remains a time-consuming proces…

cs.CV2017

Snapshot Difference Imaging using Time-of-Flight Sensors

Clara Callenberg, Felix Heide, Gordon Wetzstein +1

Computational photography encompasses a diversity of imaging techniques, but one of the core operations performed by many of them is to compute image differences. An intuitive appr…

cs.CV2022

3D GAN Inversion for Controllable Portrait Image Animation

Connor Z. Lin, David B. Lindell, Eric R. Chan +1

Millions of images of human faces are captured every single day; but these photographs portray the likeness of an individual with a fixed pose, expression, and appearance. Portrait…

cs.CV2021

Dirty Pixels: Towards End-to-End Image Processing and Perception

Steven Diamond, Vincent Sitzmann, Frank Julca-Aguilar +3

Real-world imaging systems acquire measurements that are degraded by noise, optical aberrations, and other imperfections that make image processing for human viewing and higher-lev…

cs.CV2026

Dual Ascent Diffusion for Inverse Problems

Minseo Kim, Axel Levy, Gordon Wetzstein

Ill-posed inverse problems are fundamental in many domains, ranging from astrophysics to medical imaging. Emerging diffusion models provide a powerful prior for solving these probl…

cs.CV2025

CameraCtrl: Enabling Camera Control for Text-to-Video Generation

Hao He, Yinghao Xu, Yuwei Guo +4

Controllability plays a crucial role in video generation, as it allows users to create and edit content more precisely. Existing models, however, lack control of camera pose that s…

cs.CV2023

DiffDreamer: Towards Consistent Unsupervised Single-view Scene Extrapolation with Conditional Diffusion Models

Shengqu Cai, Eric Ryan Chan, Songyou Peng +4

Scene extrapolation -- the idea of generating novel views by flying into a given image -- is a promising, yet challenging task. For each predicted frame, a joint inpainting and 3D…

cs.CV2023

PointOdyssey: A Large-Scale Synthetic Dataset for Long-Term Point Tracking

Yang Zheng, Adam W. Harley, Bokui Shen +2

We introduce PointOdyssey, a large-scale synthetic dataset, and data generation framework, for the training and evaluation of long-term fine-grained tracking algorithms. Our goal i…

cs.CV2020

State of the Art on Neural Rendering

Ayush Tewari, Ohad Fried, Justus Thies +16

Efficient rendering of photo-realistic virtual worlds is a long standing effort of computer graphics. Modern graphics techniques have succeeded in synthesizing photo-realistic imag…

cs.ET2014

A Compressive Multi-Mode Superresolution Display

Felix Heide, James Gregson, Gordon Wetzstein +2

Compressive displays are an emerging technology exploring the co-design of new optical device configurations and compressive computation. Previously, research has shown how to impr…

cs.GR2026

Garment Particles: A 2D--3D Symmetric Garment Representation for Generation and Editing

Kiyohiro Nakayama, I-Chao Shen, Ruofan Liu +3

Practical garment design spans two modes: intuitive creation from high-level intent, such as a reference image or text description, and complex low-level editing across 2D sewing p…

cs.CV2025

X-Dyna: Expressive Dynamic Human Image Animation

Di Chang, Hongyi Xu, You Xie +12

We introduce X-Dyna, a novel zero-shot, diffusion-based pipeline for animating a single human image using facial expressions and body movements derived from a driving video, that g…

cs.AI2026

MultiGen: Level-Design for Editable Multiplayer Worlds in Diffusion Game Engines

Ryan Po, David Junhao Zhang, Amir Hertz +3

Video world models have shown immense promise for interactive simulation and entertainment, but current systems still struggle with two important aspects of interactivity: user con…

cs.CV2022

CryoAI: Amortized Inference of Poses for Ab Initio Reconstruction of 3D Molecular Volumes from Real Cryo-EM Images

Axel Levy, Frédéric Poitevin, Julien Martel +6

Cryo-electron microscopy (cryo-EM) has become a tool of fundamental importance in structural biology, helping us understand the basic building blocks of life. The algorithmic chall…

cs.CV2025

BulletTime: Decoupled Control of Time and Camera Pose for Video Generation

Yiming Wang, Qihang Zhang, Shengqu Cai +7

Emerging video diffusion models achieve high visual fidelity but fundamentally couple scene dynamics with camera motion, limiting their ability to provide precise spatial and tempo…

cs.LG2025

Gaussian Mixture Flow Matching Models

Hansheng Chen, Kai Zhang, Hao Tan +5

Diffusion models approximate the denoising distribution as a Gaussian and predict its mean, whereas flow matching models reparameterize the Gaussian mean as flow velocity. However,…

cs.LG2024

Mixture of neural fields for heterogeneous reconstruction in cryo-EM

Axel Levy, Rishwanth Raghu, David Shustin +5

Cryo-electron microscopy (cryo-EM) is an experimental technique for protein structure determination that images an ensemble of macromolecules in near-physiological contexts. While…

cs.CV2024

Diffusion Self-Distillation for Zero-Shot Customized Image Generation

Shengqu Cai, Eric Chan, Yunzhi Zhang +3

Text-to-image diffusion models produce impressive results but are frustrating tools for artists who desire fine-grained control. For example, a common use case is to create images…

physics.med-ph2023

Volumetric Reconstruction Resolves Off-Resonance Artifacts in Static and Dynamic PROPELLER MRI

Annesha Ghosh, Gordon Wetzstein, Mert Pilanci +1

Off-resonance artifacts in magnetic resonance imaging (MRI) are visual distortions that occur when the actual resonant frequencies of spins within the imaging volume differ from th…

cs.CV2025

BAgger: Backwards Aggregation for Mitigating Drift in Autoregressive Video Diffusion Models

Ryan Po, Eric Ryan Chan, Changan Chen +1

Autoregressive video models are promising for world modeling via next-frame prediction, but they suffer from exposure bias: a mismatch between training on clean contexts and infere…

cs.CV2023

Instant Continual Learning of Neural Radiance Fields

Ryan Po, Zhengyang Dong, Alexander W. Bergman +1

Neural radiance fields (NeRFs) have emerged as an effective method for novel-view synthesis and 3D scene reconstruction. However, conventional training methods require access to al…

cs.CV2018

Convolutional Sparse Coding for High Dynamic Range Imaging

Ana Serrano, Felix Heide, Diego Gutierrez +2

Current HDR acquisition techniques are based on either (i) fusing multibracketed, low dynamic range (LDR) images, (ii) modifying existing hardware and capturing different exposures…

cs.CV2026

Mode Seeking meets Mean Seeking for Fast Long Video Generation

Shengqu Cai, Weili Nie, Chao Liu +8

Scaling video generation from seconds to minutes faces a critical bottleneck: while short-video data is abundant and high-fidelity, coherent long-form data is scarce and limited to…

cs.CV2022

ALTO: Alternating Latent Topologies for Implicit 3D Reconstruction

Zhen Wang, Shijie Zhou, Jeong Joon Park +5

This work introduces alternating latent topologies (ALTO) for high-fidelity reconstruction of implicit 3D surfaces from noisy point clouds. Previous work identifies that the spatia…

cs.RO2026

WorldArena 2.0: Extending Embodied World Model Benchmarking on Modality, Functionality and Platform

Yu Shang, Yinzhou Tang, Yiding Ma +22

World models have emerged as a central paradigm for embodied intelligence, enabling agents to predict action-conditioned future and reason about environmental dynamics. However, ex…

cs.CV2024

Make-it-Real: Unleashing Large Multimodal Model for Painting 3D Objects with Realistic Materials

Ye Fang, Zeyi Sun, Tong Wu +4

Physically realistic materials are pivotal in augmenting the realism of 3D assets across various applications and lighting conditions. However, existing 3D assets and generative mo…

cs.CV2026

TinyHistory: Lightweight Video History Embeddings via Two-Stage Context Learning

Lvmin Zhang, Shengqu Cai, Muyang Li +6

History context is central to autoregressive video generation, driving consistency and storytelling for both commercial models and personal use cases. For example, personal users,…

cs.CV2025

3D-Adapter: Geometry-Consistent Multi-View Diffusion for High-Quality 3D Generation

Hansheng Chen, Bokui Shen, Yulin Liu +7

Multi-view image diffusion models have significantly advanced open-domain 3D object generation. However, most existing models rely on 2D network architectures that lack inherent 3D…

cs.CV2024

Real-time 3D-aware Portrait Editing from a Single Image

Qingyan Bai, Zifan Shi, Yinghao Xu +7

This work presents 3DPE, a practical method that can efficiently edit a face image following given prompts, like reference images or text descriptions, in a 3D-aware manner. To thi…

cs.HC2024

GazeGPT: Augmenting Human Capabilities using Gaze-contingent Contextual AI for Smart Eyewear

Robert Konrad, Nitish Padmanaban, J. Gabriel Buckmaster +2

Multimodal large language models (LMMs) excel in world knowledge and problem-solving abilities. Through the use of a world-facing camera and contextual AI, emerging smart accessori…

cs.LG2022

Learning to Solve PDE-constrained Inverse Problems with Graph Networks

Qingqing Zhao, David B. Lindell, Gordon Wetzstein

Learned graph neural networks (GNNs) have recently been established as fast and accurate alternatives for principled solvers in simulating the dynamics of physical systems. In many…

cs.CV2023

Generative Neural Articulated Radiance Fields

Alexander W. Bergman, Petr Kellnhofer, Wang Yifan +3

Unsupervised learning of 3D-aware generative adversarial networks (GANs) using only collections of single-view 2D photographs has very recently made much progress. These 3D GANs, h…

cs.GR2026

MeshFlow: Mesh Generation with Equivariant Flow Matching

Qi Sun, Kiyohiro Nakayama, Jing Nathan Yan +6

Meshes are among the most common 3D scene representations, but directly generating meshes is challenging because the representation contains important symmetries, including permuta…

cs.CV2025

3DGen-Bench: Comprehensive Benchmark Suite for 3D Generative Models

Yuhan Zhang, Mengchen Zhang, Tong Wu +4

3D generation is experiencing rapid advancements, while the development of 3D evaluation has not kept pace. How to keep automatic evaluation equitably aligned with human perception…

cs.GR2018

Movie Editing and Cognitive Event Segmentation in Virtual Reality Video

Ana Serrano, Vincent Sitzmann, Jaime Ruiz-Borau +3

Traditional cinematography has relied for over a century on a well-established set of editing rules, called continuity editing, to create a sense of situational continuity. Despite…

cs.CV2023

Generative Rendering: Controllable 4D-Guided Video Generation with 2D Diffusion Models

Shengqu Cai, Duygu Ceylan, Matheus Gadelha +3

Traditional 3D content creation tools empower users to bring their imagination to life by giving them direct control over a scene's geometry, appearance, motion, and camera path. C…

cs.GR2026

Fast Wave-optics Rendering of Multiplane Images for 3D Holographic Displays

Brian Chao, Dario Seyb, Nathan Matsuda +6

Recent advances in neural rendering have unlocked unprecedented capabilities in 3D reconstruction and novel view synthesis, giving rise to applications such as virtual fly-throughs…

eess.SP2020

D-VDAMP: Denoising-based Approximate Message Passing for Compressive MRI

Christopher A. Metzler, Gordon Wetzstein

Plug and play (P&P) algorithms iteratively apply highly optimized image denoisers to impose priors and solve computational image reconstruction problems, to great effect. However,…

cs.CV2025

CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models

Qingqing Zhao, Yao Lu, Moo Jin Kim +12

Vision-language-action models (VLAs) have shown potential in leveraging pretrained vision-language models and diverse robot demonstrations for learning generalizable sensorimotor c…

eess.IV2022

MantissaCam: Learning Snapshot High-dynamic-range Imaging with Perceptually-based In-pixel Irradiance Encoding

Haley M. So, Julien N. P. Martel, Piotr Dudek +1

The ability to image high-dynamic-range (HDR) scenes is crucial in many computer vision applications. The dynamic range of conventional sensors, however, is fundamentally limited b…

cs.CV2018

Non-line-of-sight Imaging with Partial Occluders and Surface Normals

Felix Heide, Matthew O'Toole, Kai Zang +3

Imaging objects obscured by occluders is a significant challenge for many applications. A camera that could "see around corners" could help improve navigation and mapping capabilit…

q-bio.BM2022

Heterogeneous reconstruction of deformable atomic models in Cryo-EM

Youssef Nashed, Ariana Peck, Julien Martel +6

Cryogenic electron microscopy (cryo-EM) provides a unique opportunity to study the structural heterogeneity of biomolecules. Being able to explain this heterogeneity with atomic mo…

cs.CV2022

Scale-Agnostic Super-Resolution in MRI using Feature-Based Coordinate Networks

Dave Van Veen, Rogier van der Sluijs, Batu Ozturkler +9

We propose using a coordinate network decoder for the task of super-resolution in MRI. The continuous signal representation of coordinate networks enables this approach to be scale…

cs.HC2021

A Perceptual Model for Eccentricity-dependent Spatio-temporal Flicker Fusion and its Applications to Foveated Graphics

Brooke Krajancich, Petr Kellnhofer, Gordon Wetzstein

Virtual and augmented reality (VR/AR) displays strive to provide a resolution, framerate and field of view that matches the perceptual capabilities of the human visual system, all…

cs.LG2025

Solving Inverse Problems in Protein Space Using Diffusion-Based Priors

Axel Levy, Eric R. Chan, Sara Fridovich-Keil +3

The interaction of a protein with its environment can be understood and controlled via its 3D structure. Experimental methods for protein structure determination, such as X-ray cry…

cs.CV2020

MetaSDF: Meta-learning Signed Distance Functions

Vincent Sitzmann, Eric R. Chan, Richard Tucker +2

Neural implicit shape representations are an emerging paradigm that offers many potential benefits over conventional discrete representations, including memory efficiency at a high…

cs.CV2023

PixelRNN: In-pixel Recurrent Neural Networks for End-to-end-optimized Perception with Neural Sensors

Haley M. So, Laurie Bose, Piotr Dudek +1

Conventional image sensors digitize high-resolution images at fast frame rates, producing a large amount of data that needs to be transmitted off the sensor for further processing.…

cs.CV2020

Implicit Neural Representations with Periodic Activation Functions

Vincent Sitzmann, Julien N. P. Martel, Alexander W. Bergman +2

Implicitly defined, continuous, differentiable signal representations parameterized by neural networks have emerged as a powerful paradigm, offering many possible benefits over con…

cs.CV2026

VULCAN: Tool-Augmented Multi Agents for Iterative 3D Object Arrangement

Zhengfei Kuang, Rui Lin, Long Zhao +3

Despite the remarkable progress of Multimodal Large Language Models (MLLMs) in 2D vision-language tasks, their application to complex 3D scene manipulation remains underexplored. I…

cs.LG2024

Neural Control Variates with Automatic Integration

Zilu Li, Guandao Yang, Qingqing Zhao +4

This paper presents a method to leverage arbitrary neural network architecture for control variates. Control variates are crucial in reducing the variance of Monte Carlo integratio…

cs.CV2024

4D-fy: Text-to-4D Generation Using Hybrid Score Distillation Sampling

Sherwin Bahmani, Ivan Skorokhodov, Victor Rong +7

Recent breakthroughs in text-to-4D generation rely on pre-trained text-to-image and text-to-video models to generate dynamic 3D scenes. However, current text-to-4D methods face a t…

cs.CV2025

GazeFusion: Saliency-Guided Image Generation

Yunxiang Zhang, Nan Wu, Connor Z. Lin +2

Diffusion models offer unprecedented image generation power given just a text prompt. While emerging approaches for controlling diffusion models have enabled users to specify the d…

physics.app-ph2018

Sub-picosecond photon-efficient 3D imaging using single-photon sensors

Felix Heide, Steven Diamond, David B. Lindell +1

Active 3D imaging systems have broad applications across disciplines, including biological imaging, remote sensing and robotics. Applications in these domains require fast acquisit…

cs.CV2025

Video World Models with Long-term Spatial Memory

Tong Wu, Shuai Yang, Ryan Po +4

Emerging world models autoregressively generate video frames in response to actions, such as camera movements and text prompts, among other control signals. Due to limited temporal…

cs.CV2024

Streetscapes: Large-scale Consistent Street View Generation Using Autoregressive Video Diffusion

Boyang Deng, Richard Tucker, Zhengqi Li +3

We present a method for generating Streetscapes-long sequences of views through an on-the-fly synthesized city-scale scene. Our generation is conditioned by language input (e.g., c…

cs.CV2025

Textured Gaussians for Enhanced 3D Scene Appearance Modeling

Brian Chao, Hung-Yu Tseng, Lorenzo Porzi +8

3D Gaussian Splatting (3DGS) has recently emerged as a state-of-the-art 3D reconstruction and rendering technique due to its high-quality results and fast training and rendering ti…

cs.CV2025

Towards Vision-Language-Garment Models for Web Knowledge Garment Understanding and Generation

Jan Ackermann, Kiyohiro Nakayama, Guandao Yang +2

Multimodal foundation models have demonstrated strong generalization, yet their ability to transfer knowledge to specialized domains such as garment generation remains underexplore…

cs.CV2025

R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation

Meng-Hao Guo, Jiajun Xu, Yi Zhang +14

Reasoning stands as a cornerstone of intelligence, enabling the synthesis of existing knowledge to solve complex problems. Despite remarkable progress, existing reasoning benchmark…

cs.CV2023

Compositional 3D Scene Generation using Locally Conditioned Diffusion

Ryan Po, Gordon Wetzstein

Designing complex 3D scenes has been a tedious, manual process requiring domain expertise. Emerging text-to-3D generative models show great promise for making this task more intuit…

cs.CV2025

ReStyle3D: Scene-Level Appearance Transfer with Semantic Correspondences

Liyuan Zhu, Shengqu Cai, Shengyu Huang +3

We introduce ReStyle3D, a novel framework for scene-level appearance transfer from a single style image to a real-world scene represented by multiple views. The method combines exp…

cs.CV2023

Generative Novel View Synthesis with 3D-Aware Diffusion Models

Eric R. Chan, Koki Nagano, Matthew A. Chan +7

We present a diffusion-based model for 3D-aware generative novel view synthesis from as few as a single input image. Our model samples from the distribution of possible renderings…

cs.CV2021

ACORN: Adaptive Coordinate Networks for Neural Scene Representation

Julien N. P. Martel, David B. Lindell, Connor Z. Lin +3

Neural representations have emerged as a new paradigm for applications in rendering, imaging, geometric modeling, and simulation. Compared to traditional representations such as me…

cs.CV2024

Flying with Photons: Rendering Novel Views of Propagating Light

Anagh Malik, Noah Juravsky, Ryan Po +3

We present an imaging and neural rendering technique that seeks to synthesize videos of light propagating through a scene from novel, moving camera viewpoints. Our approach relies…