89 citations · 90 across the 9 of their papers we have counts for
14 papers · 1 filter
PointDiT: Pixel-Space Diffusion for Monocular Geometry Estimation
Haofei Xu, Rundi Wu, Philipp Henzler +7
State-of-the-art single-image 3D reconstruction methods often rely on complex hybrid architectures and loss functions, or compress geometry into latent spaces in order to leverage…
Pixel Cube: Diffusion-based Portrait Video Relighting Through Realistic Lighting Reproduction
Yufan Zhang, Yu Ji, Ayo Ajiboye +4
We present a diffusion-based method for relighting dynamic portrait videos with photorealism and temporal consistency. Our method is fueled by a hybrid training dataset that consis…
ViPS: Video-informed Pose Spaces for Auto-Rigged Meshes
Honglin Chen, Karran Pandey, Rundi Wu +6
Kinematic rigs provide a structured interface for articulating 3D meshes but lack any associated pose space, i.e., an explicit representation of the plausible manifold of joint con…
ZipMap: Linear-Time Stateful 3D Reconstruction via Test-Time Training
Haian Jin, Rundi Wu, Tianyuan Zhang +4
Feed-forward transformer models have driven rapid progress in 3D vision, but state-of-the-art methods such as VGGT and have a computational cost that scales quadratically wit…
VLMaterial: Procedural Material Generation with Large Vision-Language Models
Beichen Li, Rundi Wu, Armando Solar-Lezama +4
Procedural materials, represented as functional node graphs, are ubiquitous in computer graphics for photorealistic material appearance design. They allow users to perform intuitiv…
CAT4D: Create Anything in 4D with Multi-View Video Diffusion Models
Rundi Wu, Ruiqi Gao, Ben Poole +4
We present CAT4D, a method for creating 4D (dynamic 3D) scenes from monocular video. CAT4D leverages a multi-view video diffusion model trained on a diverse combination of datasets…