9 papers
PhysEditWorld: A Large-Scale Dataset Toward Physics-Editable World Models
Bin Hu, Yanwen Ma, Jiehui Huang +14
Recent game world models can synthesize visually plausible, action-conditioned rollouts. However, their interaction behaviors often remain limited to exploratory or wandering traje…
RoPE-Aware Bit Allocation for KV-Cache Quantization
Fengfeng Liang, Yuechen Zhang, Jiaya Jia
Existing low-bit KV-cache quantizers often treat each cached key as a flat vector. Under RoPE, however, a key's contribution to a future attention logit decomposes into a position-…
CoDMD: Copula-aware Distribution Matching Distillation for Fast Video Generation
Wenhu Zhang, Kun Cheng, Changyuan Wang +7
Few-step distillation for video diffusion models has attracted significant attention, driven by the urgent demand for efficient deployment in real-world scenarios. However, Distrib…
UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating
Jiehui Huang, Yuechen Zhang, Bin Xia +7
Generating a coherent multi-shot video requires structured cross-shot memory. Subject appearance, scene context, and speaker identity must persist across cuts. Existing approaches…
Training-Free Efficient Video Generation via Dynamic Token Carving
Yuechen Zhang, Jinbo Xing, Bin Xia +6
Despite the remarkable generation quality of video Diffusion Transformer (DiT) models, their practical deployment is severely hindered by extensive computational requirements. This…
Exploiting Discriminative Codebook Prior for Autoregressive Image Generation
Longxiang Tang, Ruihang Chu, Xiang Wang +6
Advanced discrete token-based autoregressive image generation systems first tokenize images into sequences of token indices with a codebook, and then model these sequences in an au…