13 papers
Anchoring Instruction Outside Mask: Exact Reference Caching for Efficient In-Context Diffusion Transformers
Yangshuai Liu, Zheming Li, Jiaao Li +4
Omnimodal generation is central to a wide range of content creation and editing applications. In-context conditioning is essential to this paradigm. It allows diffusion transformer…
Hit-and-Run Mixes as Fast as the Ball Walk
Ruizhe Zhang
Let be an isotropic convex body. We prove that the hit-and-run walk, started from any -warm distribution, reaches total-variation distance f…
Acoustic Tweezers for Magnetic Skyrmions
Chongzhou Wang, Weichao Yu
Current methods for driving magnetic skyrmions predominantly translate ensembles as a whole, lacking single-particle selectivity. Here, we propose an "acoustic tweezer" that determ…
Sekai2: From World Exploration to Interactive World Modeling
Kang He, Wenshuo Peng, Zihui Gao +3
Video world models must capture how scenes evolve over time and across viewpoints. Training them for long-horizon generation and camera control therefore benefits from long videos…
X-Stage: An Overlooked Pipeline Stage for Communication-Computation Overlap in DiT Inference
Jianwen Xian, Zhiyuan Xu, Yuchen Li +8
Fine-grained, device-initiated communication lets persistent GPU kernels in distributed diffusion transformer (DiT) inference issue remote stores and overlap data movement with Ten…
HeadCast: Casting Attention Heads for Efficient Autoregressive Video Generation
Jinliang Shen, Lianghao Su, Zheming Li +4
Autoregressive (AR) video diffusion models have become a promising paradigm for long and streaming video synthesis, but the continuously growing Key-Value (KV) cache makes attentio…