6 papers
TGT: Text-Grounded Trajectories for Locally Controlled Video Generation
Guofeng Zhang, Angtian Wang, Jacob Zhiyuan Fang +8
Text-to-video generation has advanced rapidly in visual fidelity, whereas standard methods still have limited ability to control the subject composition of generated scenes. Prior…
4D-Animal: Freely Reconstructing Animatable 3D Animals from Videos
Shanshan Zhong, Jiawei Peng, Zehan Zheng +6
Existing methods for reconstructing animatable 3D animals from videos typically rely on sparse semantic keypoints to fit parametric models. However, obtaining such keypoints is lab…
Creation-MMBench: Assessing Context-Aware Creative Intelligence in MLLM
Xinyu Fang, Zhijian Chen, Kai Lan +10
Creativity is a fundamental aspect of intelligence, involving the ability to generate novel and appropriate solutions across diverse contexts. While Large Language Models (LLMs) ha…
DINeMo: Learning Neural Mesh Models with no 3D Annotations
Weijie Guo, Guofeng Zhang, Wufei Ma +1
Category-level 3D/6D pose estimation is a crucial step towards comprehensive 3D scene understanding, which would enable a broad range of applications in robotics and embodied AI. R…
X-LRM: X-ray Large Reconstruction Model for Extremely Sparse-View Computed Tomography Recovery in One Second
Guofeng Zhang, Ruyi Zha, Hao He +4
Sparse-view 3D CT reconstruction aims to recover volumetric structures from a limited number of 2D X-ray projections. Existing feedforward methods are constrained by the scarcity o…
3DSRBench: A Comprehensive 3D Spatial Reasoning Benchmark
Wufei Ma, Haoyu Chen, Guofeng Zhang +4
3D spatial reasoning is the ability to analyze and interpret the positions, orientations, and spatial relationships of objects within the 3D space. This allows models to develop a…