9 papers
SparseCtrl-HOI: Sparse Temporal Control for Human-Object Interaction Video Generation
Shenbo Xie, Mingrui Cai, Xu Yang +2
Human-Object Interaction (HOI) video generation aims to synthesize realistic videos of humans manipulating diverse objects, serving as a promising avenue for AI-driven live streami…
Reconstruction-Anchored Diffusion Model for Text-to-Motion Generation
Yifei Liu, Changxing Ding, Ling Guo +2
Diffusion models have seen widespread adoption for text-driven human motion generation and related tasks due to their impressive generative capabilities and flexibility. However, c…
Coordinating Multiple Conditions for Trajectory-Controlled Human Motion Generation
Deli Cai, Haoyang Ma, Changxing Ding
Trajectory-controlled human motion generation aims to synthesize realistic human motions conditioned on both textual descriptions and spatial trajectories. However, existing method…
Streamlined Open-Vocabulary Human-Object Interaction Detection
Chang Sun, Dongliang Liao, Changxing Ding
Open-vocabulary human-object interaction (HOI) detection aims to localize and recognize all human-object interactions in an image, including those unseen during training. Existing…
ViHOI: Human-Object Interaction Synthesis with Visual Priors
Songjin Cai, Linjie Zhong, Ling Guo +1
Generating realistic and physically plausible 3D Human-Object Interactions (HOI) remains a key challenge in motion generation. One primary reason is that describing these physical…
Democratizing High-Fidelity Co-Speech Gesture Video Generation
Xu Yang, Shaoli Huang, Shenbo Xie +3
Co-speech gesture video generation aims to synthesize realistic, audio-aligned videos of speakers, complete with synchronized facial expressions and body gestures. This task presen…