collaborators
Showing cs.CVShow all

9 papers · 1 filter

cs.CV2026

Anchoring Instruction Outside Mask: Exact Reference Caching for Efficient In-Context Diffusion Transformers

Yangshuai Liu, Zheming Li, Jiaao Li +4

Omnimodal generation is central to a wide range of content creation and editing applications. In-context conditioning is essential to this paradigm. It allows diffusion transformer…

cs.CV2026

Sekai2: From World Exploration to Interactive World Modeling

Kang He, Wenshuo Peng, Zihui Gao +3

Video world models must capture how scenes evolve over time and across viewpoints. Training them for long-horizon generation and camera control therefore benefits from long videos…

cs.CV2026

HeadCast: Casting Attention Heads for Efficient Autoregressive Video Generation

Jinliang Shen, Lianghao Su, Zheming Li +4

Autoregressive (AR) video diffusion models have become a promising paradigm for long and streaming video synthesis, but the continuously growing Key-Value (KV) cache makes attentio…

cs.CV2026

AlayaWorld: Long-Horizon and Playable Video World Generation

AlayaWorld Team, Kaipeng Zhang, Chuanhao Li +14

Game worlds have traditionally been built through labor-intensive production pipelines, making them costly to develop, difficult to customization, and expensive to modify after dep…

cs.CV2026

ScalingAttention: Discovering Intrinsic Sparse Attention Topology for Video Diffusion Transformers

Ruiliang Zhou, Xuecheng Wu, Kang He +6

While Diffusion Transformers (DiTs) have revolutionized high-fidelity video generation, their reliance on 3D full attention creates a quadratic computational bottleneck. Existing s…

cs.CV2026

Qwen-Image-2.0 Technical Report

Bing Zhao, Chenfei Wu, Deqing Li +72

We present Qwen-Image-2.0, an omni-capable image generation foundation model that unifies high-fidelity generation and precise image editing within a single framework. Despite rece…