12 papers
NeuROK: Generative 4D Neural Object Kinematics
Chen Geng, Guangzhao He, Yue Gao +3
Data-driven approaches have revolutionized 3D vision, enabling transformers to effectively reconstruct and generate static 3D objects. However, generating simulative 4D dynamics --…
Thinking with Spatial Code for Physical-World Video Reasoning
Jieneng Chen, Wenxin Ma, Ruisheng Yuan +3
We introduce Thinking with Spatial Code, a framework that transforms RGB video into explicit, temporally coherent 3D representations for physical-world visual question answering. W…
Choreographing a World of Dynamic Objects
Yanzhe Lyu, Chen Geng, Karthik Dharmarajan +4
Dynamic objects in our physical 4D (3D + time) world are constantly evolving, deforming, and interacting with other objects, leading to diverse 4D scene dynamics. In this paper, we…
Model-Based Policy Adaptation for Closed-Loop End-to-End Autonomous Driving
Haohong Lin, Yunzhi Zhang, Wenhao Ding +2
End-to-end (E2E) autonomous driving models have demonstrated strong performance in open-loop evaluations but often suffer from cascading errors and poor generalization in closed-lo…
Coupled Diffusion Sampling for Training-Free Multi-View Image Editing
Hadi Alzayer, Yunzhi Zhang, Chen Geng +2
We present an inference-time diffusion sampling method to perform multi-view consistent image editing using pre-trained 2D image editing models. These models can independently prod…
Ctrl-VI: Controllable Video Synthesis via Variational Inference
Haoyi Duan, Yunzhi Zhang, Yilun Du +1
Many video workflows benefit from a mixture of user controls with varying granularity, from exact 4D object trajectories and camera paths to coarse text prompts, while existing vid…