collaborators

6 papers

cs.CV2025

TGT: Text-Grounded Trajectories for Locally Controlled Video Generation

Guofeng Zhang, Angtian Wang, Jacob Zhiyuan Fang +8

Text-to-video generation has advanced rapidly in visual fidelity, whereas standard methods still have limited ability to control the subject composition of generated scenes. Prior…

cs.CV2025

4D-Animal: Freely Reconstructing Animatable 3D Animals from Videos

Shanshan Zhong, Jiawei Peng, Zehan Zheng +6

Existing methods for reconstructing animatable 3D animals from videos typically rely on sparse semantic keypoints to fit parametric models. However, obtaining such keypoints is lab…

cs.CV2025

Creation-MMBench: Assessing Context-Aware Creative Intelligence in MLLM

Xinyu Fang, Zhijian Chen, Kai Lan +10

Creativity is a fundamental aspect of intelligence, involving the ability to generate novel and appropriate solutions across diverse contexts. While Large Language Models (LLMs) ha…

cs.CV2025

DINeMo: Learning Neural Mesh Models with no 3D Annotations

Weijie Guo, Guofeng Zhang, Wufei Ma +1

Category-level 3D/6D pose estimation is a crucial step towards comprehensive 3D scene understanding, which would enable a broad range of applications in robotics and embodied AI. R…

eess.IV2025

X-LRM: X-ray Large Reconstruction Model for Extremely Sparse-View Computed Tomography Recovery in One Second

Guofeng Zhang, Ruyi Zha, Hao He +4

Sparse-view 3D CT reconstruction aims to recover volumetric structures from a limited number of 2D X-ray projections. Existing feedforward methods are constrained by the scarcity o…

cs.CV2024

3DSRBench: A Comprehensive 3D Spatial Reasoning Benchmark

Wufei Ma, Haoyu Chen, Guofeng Zhang +4

3D spatial reasoning is the ability to analyze and interpret the positions, orientations, and spatial relationships of objects within the 3D space. This allows models to develop a…