6 papers
Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO
Yunze Tong, Mushui Liu, Canyu Zhao +7
Deploying GRPO on Flow Matching models has proven effective for text-to-image generation. However, existing paradigms typically propagate an outcome-based reward to all preceding d…
ProFashion: Prototype-guided Fashion Video Generation with Multiple Reference Images
Xianghao Kong, Qiaosong Qi, Yuanbin Wang +3
Fashion video generation aims to synthesize temporally consistent videos from reference images of a designated character. Despite significant progress, existing diffusion-based met…
Does Hearing Help Seeing? Investigating Audio-Video Joint Denoising for Video Generation
Jianzong Wu, Hao Lian, Dachao Hao +5
Recent audio-video generative systems suggest that coupling modalities benefits not only audio-video synchrony but also the video modality itself. We pose a fundamental question: D…
SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers
Zhengcong Fei, Hao Jiang, Di Qiu +8
The generation and editing of audio-conditioned talking portraits guided by multimodal inputs, including text, images, and videos, remains under explored. In this paper, we present…
SmoothCache: A Universal Inference Acceleration Technique for Diffusion Transformers
Joseph Liu, Joshua Geddes, Ziyu Guo +2
Diffusion Transformers (DiT) have emerged as powerful generative models for various tasks, including image, video, and speech synthesis. However, their inference process remains co…
Unsupervised Cross-Domain Regression for Fine-grained 3D Game Character Reconstruction
Qi Wen, Xiang Wen, Hao Jiang +5
With the rise of the ``metaverse'' and the rapid development of games, it has become more and more critical to reconstruct characters in the virtual world faithfully. The immersive…