3 papers
cs.CV2026
SparseCtrl-HOI: Sparse Temporal Control for Human-Object Interaction Video Generation
Shenbo Xie, Mingrui Cai, Xu Yang +2
Human-Object Interaction (HOI) video generation aims to synthesize realistic videos of humans manipulating diverse objects, serving as a promising avenue for AI-driven live streami…
cs.CV2026
A Creative Agent is Worth a 64-Token Template
Ruixiao Shi, Fu Feng, Yucheng Xie +3
Text-to-image (T2I) models have substantially improved image fidelity and prompt adherence, yet their creativity remains constrained by reliance on discrete natural language prompt…
cs.CV2025
Democratizing High-Fidelity Co-Speech Gesture Video Generation
Xu Yang, Shaoli Huang, Shenbo Xie +3
Co-speech gesture video generation aims to synthesize realistic, audio-aligned videos of speakers, complete with synchronized facial expressions and body gestures. This task presen…