6 papers
Soul: Breathe Life into Digital Human for High-fidelity Long-term Multimodal Animation
Jiangning Zhang, Junwei Zhu, Zhenye Gan +14
We propose a multimodal-driven framework for high-fidelity long-term digital human animation termed , which generates semantically coherent videos from a single-fram…
Transform Trained Transformer: Accelerating Naive 4K Video Generation Over 10
Jiangning Zhang, Junwei Zhu, Teng Hu +7
Native 4K (21603840) video generation remains a critical challenge due to the quadratic computational explosion of full-attention as spatiotemporal resolution increases, ma…
SwiftVideo: A Unified Framework for Few-Step Video Generation through Trajectory-Distribution Alignment
Yanxiao Sun, Jiafu Wu, Yun Cao +6
Diffusion-based or flow-based models have achieved significant progress in video synthesis but require multiple iterative sampling steps, which incurs substantial computational ove…
StrandDesigner: Towards Practical Strand Generation with Sketch Guidance
Na Zhang, Moran Li, Chengming Xu +6
Realistic hair strand generation is crucial for applications like computer graphics and virtual reality. While diffusion models can generate hairstyles from text or images, these i…
Collaborative Face Experts Fusion in Video Generation: Boosting Identity Consistency Across Large Face Poses
Yuji Wang, Moran Li, Xiaobin Hu +7
Current video generation models struggle with identity preservation under large face poses, primarily facing two challenges: the difficulty in exploring an effective mechanism to i…
Identity-Preserving Text-to-Video Generation Guided by Simple yet Effective Spatial-Temporal Decoupled Representations
Yuji Wang, Moran Li, Xiaobin Hu +7
Identity-preserving text-to-video (IPT2V) generation, which aims to create high-fidelity videos with consistent human identity, has become crucial for downstream applications. Howe…