17 papers
FVAttn: Adaptive Sparse Attention with Runtime Load Balancing for Video Generation
Hao Liu, Chenghuan Huang, Ye Huang +6
Video Diffusion Transformers process long spatio-temporal sequences, making self-attention the main bottleneck in high-resolution video generation. Training-free sparse attention r…
Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing
Feng Wang, Canmiao Fu, Zhipeng Huang +3
Recent unified multimodal models show a single architecture can jointly perform vision/language understanding and image generation/editing. However, they repeatedly feed all histor…
EMOSH: Expressive Motion and Shape Disentanglement for Human Animation
Dongbin Zhang, Hao Liu, Binquan Dai +5
High-fidelity and expressive controllable human animation is essential for content creation and digital avatar applications. However, existing methods face a dilemma between expres…
PortraitGen: Exemplar-Driven GRPO with Dual-Reward Guidance for Photorealistic Portrait Generation
Xiaomin Li, Qian Liang, Yinan Li +5
Reinforcement Learning like Group Relative Policy Optimization (GRPO) has significantly advanced text-to-image post-training. However, current methods often favor superficial aesth…
DRM: Diffusion-based Reward Model With Step-wise Guidance
Jaxon Zhang, Binxin Yang, Hubery Yin +2
Current mainstream methods of aligning diffusion models with human preferences typically employ VLM-based reward models. However, these reward models, pre-trained for semantic alig…
Chorus II: Cross-Request Sparsity Reuse for Efficient Image-to-Video Generation
Hao Liu, Chenghuan Huang, Xing Cai +5
Serving diffusion models for image-to-video generation is computationally expensive, posing significant challenges for large-scale deployment. Real I2V workloads often contain simi…