From the 1 of 14 linked papers with an AI index.
14 papers
UniMoCa: Unifying Motion and Camera Controls as Visual Proxies for Faithful Human Video Generation
Liming Tan, Ye Chen, Hao Zhang +3
Controlling human motion and camera movement is essential for faithful human-oriented video generation, yet remains challenging in multi-person scenes with large body motions, occl…
RegHead: Non-Humanoid Head Blendshapes via Feed-Forward Registration
Jiahao Luo, Hao Zhang, Jianqi Chen +9
RegHead is a framework that builds semantic blendshape sets for animatable non‑humanoid head avatars using a fast feed‑forward registration model and a large dataset of shared expr…
Flex4DHuman: Flexible Multi-view Video Diffusion for 4D Human Reconstruction
Jen-Hao Cheng, Yipeng Wang, Hao Zhang +2
We present Flex4DHuman, a multi-view video diffusion model that transforms a monocular or sparse multi-view video of a dynamic subject into synchronized dense multi-view videos usi…
World Tracing: Generative Pixel-Aligned Geometry Beyond the Visible
Hao Zhang, Mohamed El Banani, Jen-Hao Cheng +6
Image-to-3D methods often trade off faithfulness and completeness: depth estimators are anchored to input pixels but stop at the visible surface, while image-to-3D models generate…
VL-DINO: Leveraging CLIP Vision-Language Knowledge for Open-Vocabulary Object Detectio
Hao Zhang, Qinran Lin, Linqi Song +1
Vision-language models like CLIP can provide rich semantic priors for open-vocabulary object detection. However, jointly integrating both textual and visual knowledge into detectio…
Echo4DIR: 4D Implicit Heart Reconstruction from 2D Echocardiography Videos
Yanan Liu, Qinya Li, Hao Zhang +5
Reconstructing 4D (3D+t) cardiac geometry from sparse 2D echocardiography is highly desirable yet fundamentally challenged by geometric ambiguity and temporal discontinuity. To tac…