11 papers
RGB-D Video Generation for Improving Human-to-Robot Object Handover Prediction
Tianyu Sun, Zhoujie Fu, Zihui Gao +2
Human-to-robot (H2R) object handover is a fundamental capability for human-robot collaboration, yet progress is hindered by the scarcity of large-scale, human-centric datasets and…
VLM-Guided Group Preference Alignment for Diffusion-based Human Mesh Recovery
Wenhao Shen, Hao Wang, Wanqi Yin +5
Human mesh recovery (HMR) from a single RGB image is inherently ambiguous, as multiple 3D poses can correspond to the same 2D observation. Recent diffusion-based methods tackle thi…
MVAnimate: Enhancing Character Animation with Multi-View Optimization
Tianyu Sun, Zhoujie Fu, Bang Zhang +1
The demand for realistic and versatile character animation has surged, driven by its wide-ranging applications in various domains. However, the animation generation algorithms mode…
Bridging Your Imagination with Audio-Video Generation via a Unified Director
Jiaxu Zhang, Tianshu Hu, Yuan Zhang +4
Existing AI-driven video creation systems typically treat script drafting and key-shot design as two disjoint tasks: the former relies on large language models, while the latter de…
ViSA: 3D-Aware Video Shading for Real-Time Upper-Body Avatar Creation
Fan Yang, Heyuan Li, Peihao Li +9
Generating high-fidelity upper-body 3D avatars from one-shot input image remains a significant challenge. Current 3D avatar generation methods, which rely on large reconstruction m…
InteractPro: A Unified Framework for Motion-Aware Image Composition
Weijing Tao, Xiaofeng Yang, Miaomiao Cui +1
We introduce InteractPro, a comprehensive framework for dynamic motion-aware image composition. At its core is InteractPlan, an intelligent planner that leverages a Large Vision La…