5 papers
HEIR: Learning Graph-Based Motion Hierarchies
Cheng Zheng, William Koch, Baiang Li +1
Hierarchical structures of motion exist across research fields, including computer vision, graphics, and robotics, where complex dynamics typically arise from coordinated interacti…
Can Video Diffusion Model Reconstruct 4D Geometry?
Jinjie Mai, Wenxuan Zhu, Haozhe Liu +4
Reconstructing dynamic 3D scenes (i.e., 4D geometry) from monocular video is an important yet challenging problem. Conventional multiview geometry-based approaches often struggle w…
4D-Bench: Benchmarking Multi-modal Large Language Models for 4D Object Understanding
Wenxuan Zhu, Bing Li, Cheng Zheng +8
Multimodal Large Language Models (MLLMs) have demonstrated impressive 2D image/video understanding capabilities. However, there are no publicly standardized benchmarks to assess th…
MovieCharacter: A Tuning-Free Framework for Controllable Character Video Synthesis
Di Qiu, Zheng Chen, Rui Wang +4
Recent advancements in character video synthesis still depend on extensive fine-tuning or complex 3D modeling processes, which can restrict accessibility and hinder real-time appli…
Vivid-ZOO: Multi-View Video Generation with Diffusion Model
Bing Li, Cheng Zheng, Wenxuan Zhu +4
While diffusion models have shown impressive performance in 2D image/video generation, diffusion-based Text-to-Multi-view-Video (T2MVid) generation remains underexplored. The new c…