6 papers
VMonarch: Efficient Video Diffusion Transformers with Structured Attention
Cheng Liang, Haoxian Chen, Liang Hou +4
The quadratic complexity of the attention mechanism severely limits the context scalability of Video Diffusion Transformers (DiTs). We find that the highly sparse spatio-temporal a…
Video-as-Answer: Predict and Generate Next Video Event with Joint-GRPO
Junhao Cheng, Liang Hou, Xin Tao +1
While language models have become impactful in many real-world applications, video generation remains largely confined to entertainment. Motivated by video's inherent capacity to d…
Score Augmentation for Diffusion Models
Liang Hou, Yuan Gao, Boyuan Jiang +6
Diffusion models have achieved remarkable success in generative modeling. However, this study confirms the existence of overfitting in diffusion model training, particularly in dat…
VMoBA: Mixture-of-Block Attention for Video Diffusion Models
Jianzong Wu, Liang Hou, Haotian Yang +5
The quadratic complexity of full attention mechanisms poses a significant bottleneck for Video Diffusion Models (VDMs) aiming to generate long-duration, high-resolution videos. Whi…
MTV-Inpaint: Multi-Task Long Video Inpainting
Shiyuan Yang, Zheng Gu, Liang Hou +4
Video inpainting involves modifying local regions within a video, ensuring spatial and temporal consistency. Most existing methods focus primarily on scene completion (i.e., fillin…
Boosting Resolution Generalization of Diffusion Transformers with Randomized Positional Encodings
Liang Hou, Cong Liu, Mingwu Zheng +4
Resolution generalization in image generation tasks enables the production of higher-resolution images with lower training resolution overhead. However, a key obstacle for diffusio…