5 papers · 1 filter
Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers
Chongjian Ge, Hanwen Jiang, Tianyu Wang +9
Visual generation increasingly requires high-resolution images, long videos, and multimodal context, making the quadratic cost of full attention prohibitive. We introduce Chimera,…
Rethinking Training Dynamics in Scale-wise Autoregressive Generation
Gengze Zhou, Chongjian Ge, Hao Tan +2
Recent advances in autoregressive (AR) generative models have produced increasingly powerful systems for media synthesis. Among them, next-scale prediction has emerged as a popular…
RELIC: Interactive Video World Model with Long-Horizon Memory
Yicong Hong, Yiqun Mei, Chongjian Ge +11
A truly interactive world model requires three key ingredients: real-time long-horizon streaming, consistent spatial memory, and precise user control. However, most existing approa…
VEGGIE: Instructional Editing and Reasoning Video Concepts with Grounded Generation
Shoubin Yu, Difan Liu, Ziqiao Ma +5
Recent video diffusion models have enhanced video editing, but it remains challenging to handle instructional editing and diverse tasks (e.g., adding, removing, changing) within a…
Progressive Autoregressive Video Diffusion Models
Desai Xie, Zhan Xu, Yicong Hong +5
Current frontier video diffusion models have demonstrated remarkable results at generating high-quality videos. However, they can only generate short video clips, normally around 1…