8 papers
Incantation: Natural Language as the Action Interface for Multi-Entity Video World Models
Shangwen Zhu, Qianyu Peng, Zhao Pu +12
Modern interactive video world models have achieved impressive visual fidelity, yet lack fine-grained multi-entity control and cross-entity, cross-world generalization. We trace th…
DySink: Dynamic Frame Sinks for Autoregressive Long Video Generation
Bo Ye, Xinyu Cui, Jian Zhao +2
Autoregressive long video generation often adopts bounded-memory streaming for efficiency, typically combining local windows for short-term continuity with static early-frame sinks…
SCOPE: Simulating Cross-game Operations in Playable Environments for FPS World Models
Zizhao Tong, Yeying Jin, Hongfeng Lai +11
Interactive world models for first-person shooter (FPS) games must resolve high-frequency overlapping control signals at every frame without disrupting unaffected regions. Existing…
Moment-Reenacting: Inverse Motion Degradation with Cross-shutter Guidance
Xiang Ji, Guixu Lin, Zhengwei Yin +2
Motion degradation, manifested as blur in global shutter (GS) images or rolling shutter (RS) distortion in RS counterparts, remains a fundamental challenge in computational imaging…
All-in-One Transferring Image Compression from Human Perception to Multi-Machine Perception
Jiancheng Zhao, Xiang Ji, Yinqiang Zheng
Efficiently transferring Learned Image Compression (LIC) model from human perception to machine perception is an emerging challenge in vision-centric representation learning. Exist…
AniCrafter: Customizing Realistic Human-Centric Animation via Avatar-Background Conditioning in Video Diffusion Models
Muyao Niu, Mingdeng Cao, Yifan Zhan +8
Recent advances in video diffusion models have substantially enhanced character animation techniques. However, existing methods primarily depend on structural conditions, such as D…