4 papers
UCM: Unified Modeling of Camera Control and Memory with Time-aware Positional Encoding Warping for World Models
Tianxing Xu, Zixuan Wang, Guangyuan Wang +5
World models based on video generation demonstrate remarkable potential for simulating interactive environments yet suffer from persistent difficulties in two key areas: maintainin…
PortraitDirector: A Hierarchical Disentanglement Framework for Controllable and Real-time Facial Reenactment
Chaonan Ji, Jinwei Qi, Sheng Xu +2
Existing facial reenactment methods struggle with a trade-off between expressiveness and fine-grained controllability. Holistic facial reenactment models often sacrifice granular c…
Co-speech Gesture Video Generation via Motion-Based Graph Retrieval
Yafei Song, Peng Zhang, Bang Zhang
Synthesizing synchronized and natural co-speech gesture videos remains a formidable challenge. Recent approaches have leveraged motion graphs to harness the potential of existing v…
MotionRAG-Diff: A Retrieval-Augmented Diffusion Framework for Long-Term Music-to-Dance Generation
Mingyang Huang, Peng Zhang, Bang Zhang
Generating long-term, coherent, and realistic music-conditioned dance sequences remains a challenging task in human motion synthesis. Existing approaches exhibit critical limitatio…