5 papers
MoCA: Identity-Preserving Text-to-Video Generation via Mixture of Cross Attention
Qi Xie, Yongjia Ma, Donglin Di +2
Achieving ID-preserving text-to-video (T2V) generation remains challenging despite recent advances in diffusion-based models. Existing approaches often fail to capture fine-grained…
MAGE:A Multi-stage Avatar Generator with Sparse Observations
Fangyu Du, Yang Yang, Xuehao Gao +1
Inferring full-body poses from Head Mounted Devices, which capture only 3-joint observations from the head and wrists, is a challenging task with wide AR/VR applications. Previous…
Jointly Understand Your Command and Intention:Reciprocal Co-Evolution between Scene-Aware 3D Human Motion Synthesis and Analysis
Xuehao Gao, Yang Yang, Shaoyi Du +2
As two intimate reciprocal tasks, scene-aware human motion synthesis and analysis require a joint understanding between multiple modalities, including 3D body motions, 3D scenes, a…
EigenActor: Variant Body-Object Interaction Generation Evolved from Invariant Action Basis Reasoning
Xuehao Gao, Yang Yang, Shaoyi Du +3
This paper explores a cross-modality synthesis task that infers 3D human-object interactions (HOIs) from a given text-based instruction. Existing text-to-HOI synthesis methods main…
Multi-Condition Latent Diffusion Network for Scene-Aware Neural Human Motion Prediction
Xuehao Gao, Yang Yang, Yang Wu +2
Inferring 3D human motion is fundamental in many applications, including understanding human activity and analyzing one's intention. While many fruitful efforts have been made to h…