9 papers · 1 filter
Anchoring Instruction Outside Mask: Exact Reference Caching for Efficient In-Context Diffusion Transformers
Yangshuai Liu, Zheming Li, Jiaao Li +4
Omnimodal generation is central to a wide range of content creation and editing applications. In-context conditioning is essential to this paradigm. It allows diffusion transformer…
Sekai2: From World Exploration to Interactive World Modeling
Kang He, Wenshuo Peng, Zihui Gao +3
Video world models must capture how scenes evolve over time and across viewpoints. Training them for long-horizon generation and camera control therefore benefits from long videos…
HeadCast: Casting Attention Heads for Efficient Autoregressive Video Generation
Jinliang Shen, Lianghao Su, Zheming Li +4
Autoregressive (AR) video diffusion models have become a promising paradigm for long and streaming video synthesis, but the continuously growing Key-Value (KV) cache makes attentio…
AlayaWorld: Long-Horizon and Playable Video World Generation
AlayaWorld Team, Kaipeng Zhang, Chuanhao Li +14
Game worlds have traditionally been built through labor-intensive production pipelines, making them costly to develop, difficult to customization, and expensive to modify after dep…
ScalingAttention: Discovering Intrinsic Sparse Attention Topology for Video Diffusion Transformers
Ruiliang Zhou, Xuecheng Wu, Kang He +6
While Diffusion Transformers (DiTs) have revolutionized high-fidelity video generation, their reliance on 3D full attention creates a quadratic computational bottleneck. Existing s…
Qwen-Image-2.0 Technical Report
Bing Zhao, Chenfei Wu, Deqing Li +72
We present Qwen-Image-2.0, an omni-capable image generation foundation model that unifies high-fidelity generation and precise image editing within a single framework. Despite rece…