From the 1 of 14 linked papers with an AI index.
12 papers · 1 filter
Video Models as Native 4D Renderers: World-Grounded Conditioning from Animated Mesh
Junhao Chen, Mingjin Chen, Henghaofan Zhang +10
Pretrained video diffusion models can act as renderers when the desired scene state is already specified by an animated mesh, a camera trajectory, and a reference image. This 4D ge…
ViP-Rig: Visual-Prompted Controllable Rigging
Zihan Qin, Mingze Sun, Yifan Mao +6
ViP-Rig is a framework that lets users control 3D rigging by providing 2D visual prompts, enabling both initial skeleton creation and later editing of skeletons and skinning weight…
4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding
Zhangquan Chen, Manyuan Zhang, Xinlei Yu +9
Dynamic spatial reasoning from monocular video is essential for bridging visual intelligence and the physical world, yet remains challenging for vision-language models (VLMs). Prio…
NFR: Neural Feature-Guided Non-Rigid Shape Registration
Zhangquan Chen, Puhua Jiang, Mingze Sun +1
In this paper, we propose a novel learning-based framework for 3D shape registration, which overcomes the challenges of significant non-rigid deformation and partiality undergoing…
Think with 3D: Geometric Imagination Grounded Spatial Reasoning from Limited Views
Zhangquan Chen, Manyuan Zhang, Xinlei Yu +8
Though recent advances in vision-language models (VLMs) have achieved remarkable progress across a wide range of multimodal tasks, understanding 3D spatial relationships from limit…
SIFThinker: Spatially-Aware Image Focus for Visual Reasoning
Zhangquan Chen, Ruihui Zhao, Chuwei Luo +4
Current multimodal large language models (MLLMs) still face significant challenges in complex visual tasks (e.g., spatial understanding, fine-grained perception). Prior methods hav…