10 papers
ForeHOI: Feed-forward 3D Object Reconstruction from Daily Hand-Object Interaction Videos
Yuantao Chen, Jiahao Chang, Chongjie Ye +4
The ubiquity of monocular videos capturing daily hand-object interactions presents a valuable resource for embodied intelligence. While 3D hand reconstruction from in-the-wild vide…
EI-Part: Explode for Completion and Implode for Refinement
Wanhu Sun, Zhongjin Luo, Heliang Zheng +6
Part-level 3D generation is crucial for various downstream applications, including gaming, film production, and industrial design. However, decomposing a 3D shape into geometricall…
MLLM-4D: Towards Visual-based Spatial-Temporal Intelligence
Xingyilang Yin, Chengzhengxu Li, Jiahao Chang +2
Humans are born with vision-based 4D spatial-temporal intelligence, which enables us to perceive and reason about the evolution of 3D space over time from purely visual inputs. Des…
ReconViaGen: Towards Accurate Multi-view 3D Object Reconstruction via Generation
Jiahao Chang, Chongjie Ye, Yushuang Wu +6
Existing multi-view 3D object reconstruction methods heavily rely on sufficient overlap between input views, where occlusions and sparse coverage in practice frequently yield sever…
MV-Performer: Taming Video Diffusion Model for Faithful and Synchronized Multi-view Performer Synthesis
Yihao Zhi, Chenghong Li, Hongjie Liao +6
Recent breakthroughs in video generation, powered by large-scale datasets and diffusion techniques, have shown that video diffusion models can function as implicit 4D novel view sy…
VC-Agent: An Interactive Agent for Customized Video Dataset Collection
Yidan Zhang, Mutian Xu, Yiming Hao +6
Facing scaling laws, video data from the internet becomes increasingly important. However, collecting extensive videos that meet specific needs is extremely labor-intensive and tim…